HamburgerMenu
hirist

Job Description

Role Overview :

The AI Platform Operations Engineer is part of the 24x7 operations team supporting the Central AI Kitchen platform and its AI services. The role is responsible for continuous monitoring, first-level incident detection and troubleshooting, execution of approved recovery procedures, and timely escalation to the appropriate platform, infrastructure and application support teams.

Tech Stack :

- Red Hat OpenShift

- Kubernetes

- Linux

- GitOps

Key Responsibilities :

24x7 Platform Monitoring :

- Monitor platform and application dashboards, alerts and operational mailboxes.

- Identify availability, infrastructure, application and service degradation events.

- Acknowledge alerts promptly and determine initial severity based on established procedures.

- Maintain accurate shift handover and operational records.

First-Level Incident Troubleshooting :

- Perform basic network connectivity checks including ping, DNS resolution and endpoint connectivity.

- Check OpenShift/Kubernetes resource status including nodes, pods, deployments and services.

- Review basic application and platform logs to identify common failure conditions.

- Check infrastructure and service health using approved dashboards and operational tools.

- Collect relevant diagnostic information before escalation.

Incident Recovery :

- Execute documented recovery procedures and operational runbooks.

- Restart or redeploy affected workloads using approved GitOps processes.

- Verify service recovery through dashboards, health checks and application endpoints.

- Escalate when recovery procedures are unsuccessful or when an incident falls outside the approved operating scope.

Incident Coordination & Escalation :

- Create and maintain incident tickets with accurate timestamps, symptoms, actions and observations.

- Engage the appropriate application, platform, infrastructure or network teams based on established escalation procedures.

- Provide clear status updates during active incidents.

- Support incident bridges by providing operational information and executing actions requested by L2/L3 engineers.

- Ensure effective handover of unresolved incidents between shifts.

Operational Procedures :

- Follow established Standard Operating Procedures (SOPs), runbooks and change-management processes.

- Document newly encountered symptoms and successful troubleshooting steps.

- Highlight recurring alerts or operational problems to senior platform engineers.

- Participate in operational drills and recovery exercises.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...