Posted on: 17/09/2026
Key Responsibilities :
24x7 Platform Monitoring :
- Monitor platform and application dashboards, alerts and operational mailboxes.
- Identify availability, infrastructure, application and service degradation events.
- Acknowledge alerts promptly and determine initial severity based on established procedures.
- Maintain accurate shift handover and operational records.
First-Level Incident Troubleshooting :
- Perform basic network connectivity checks including ping, DNS resolution and endpoint connectivity.
- Check OpenShift/Kubernetes resource status including nodes, pods, deployments and services.
- Review basic application and platform logs to identify common failure conditions.
- Check infrastructure and service health using approved dashboards and operational tools.
- Collect relevant diagnostic information before escalation.
Incident Recovery :
- Execute documented recovery procedures and operational runbooks.
- Restart or redeploy affected workloads using approved GitOps processes.
- Verify service recovery through dashboards, health checks and application endpoints.
- Escalate when recovery procedures are unsuccessful or when an incident falls outside the approved operating scope.
Incident Coordination & Escalation :
- Create and maintain incident tickets with accurate timestamps, symptoms, actions and observations.
- Engage the appropriate application, platform, infrastructure or network teams based on established escalation procedures.
- Provide clear status updates during active incidents.
- Support incident bridges by providing operational information and executing actions requested by L2/L3 engineers.
- Ensure effective handover of unresolved incidents between shifts.
Operational Procedures :
- Follow established Standard Operating Procedures (SOPs), runbooks and change-management processes.
- Document newly encountered symptoms and successful troubleshooting steps.
- Highlight recurring alerts or operational problems to senior platform engineers.
- Participate in operational drills and recovery exercises.
Scope of Authority :
Engineer may independently :
- Acknowledge and investigate alerts.
- Perform approved diagnostic commands and health checks.
- Execute documented L1 recovery procedures.
- Trigger approved GitOps redeployments.
- Open incidents and engage predefined support teams.
- Escalate incidents based on severity and runbook criteria.
Engineer must escalate :
- Changes requiring manual modification of production infrastructure.
- Unauthorised configuration or source-code changes.
- Security incidents or suspected compromises.
- Infrastructure failures requiring RE intervention.
- Incidents where documented recovery procedures fail.
- Major incidents requiring business or management decisions.
Qualifications & Experience :
- Bachelor's degree in Computer Science, Information Technology, Engineering or related discipline.
- 2 years of IT operations, infrastructure, cloud or application support experience.
- Basic Linux command-line knowledge.
- Basic networking knowledge including IP addressing, DNS, ports and connectivity troubleshooting.
- Basic understanding of containers and Kubernetes concepts.
- Ability to follow technical procedures accurately.
- Good written and verbal communication skills.
- Willingness and ability to work in a 24x7 shift environment.
Good to Have :
- Exposure to Kubernetes or Red Hat OpenShift.
- Exposure to Git and GitOps concepts.
- Familiarity with monitoring tools such as Grafana and Prometheus.
- Basic scripting experience with Bash or Python.
- Familiarity with incident-management or ITIL processes.
- Exposure to cloud or data-centre infrastructure.
- Interest in AI/ML infrastructure and GPU-based platforms.
Key Competencies :
- Systematic troubleshooting.
- Attention to detail.
- Ability to remain structured during incidents.
- Clear communication and escalation.
- Discipline in following operational procedures.
- Willingness to learn.
- Teamwork across geographically distributed support teams.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
DevOps / Cloud
Job Code
1672153