Posted on: 30/09/2026
Role Overview :
The AI Platform Operations Engineer is part of the 24x7 operations team supporting the Central AI Kitchen platform and its AI services. The role is responsible for continuous monitoring, first-level incident detection and troubleshooting, execution of approved recovery procedures, and timely escalation to the appropriate platform, infrastructure and application support teams.
Tech Stack :
- Red Hat OpenShift
- Kubernetes
- Linux
- GitOps
Key Responsibilities :
24x7 Platform Monitoring :
- Monitor platform and application dashboards, alerts and operational mailboxes.
- Identify availability, infrastructure, application and service degradation events.
- Acknowledge alerts promptly and determine initial severity based on established procedures.
- Maintain accurate shift handover and operational records.
First-Level Incident Troubleshooting :
- Perform basic network connectivity checks including ping, DNS resolution and endpoint connectivity.
- Check OpenShift/Kubernetes resource status including nodes, pods, deployments and services.
- Review basic application and platform logs to identify common failure conditions.
- Check infrastructure and service health using approved dashboards and operational tools.
- Collect relevant diagnostic information before escalation.
Incident Recovery :
- Execute documented recovery procedures and operational runbooks.
- Restart or redeploy affected workloads using approved GitOps processes.
- Verify service recovery through dashboards, health checks and application endpoints.
- Escalate when recovery procedures are unsuccessful or when an incident falls outside the approved operating scope.
Incident Coordination & Escalation :
- Create and maintain incident tickets with accurate timestamps, symptoms, actions and observations.
- Engage the appropriate application, platform, infrastructure or network teams based on established escalation procedures.
- Provide clear status updates during active incidents.
- Support incident bridges by providing operational information and executing actions requested by L2/L3 engineers.
- Ensure effective handover of unresolved incidents between shifts.
Operational Procedures :
- Follow established Standard Operating Procedures (SOPs), runbooks and change-management processes.
- Document newly encountered symptoms and successful troubleshooting steps.
- Highlight recurring alerts or operational problems to senior platform engineers.
- Participate in operational drills and recovery exercises.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Systems Administration
Job Code
1675781