HamburgerMenu
hirist

L1 Platform Engineer - Cloud Infrastructure

Anvaya Info Solutions
2 - 5 Years
Pune

Posted on: 17/09/2026

Job Description

Key Responsibilities :

24x7 Platform Monitoring :

- Monitor platform and application dashboards, alerts and operational mailboxes.

- Identify availability, infrastructure, application and service degradation events.

- Acknowledge alerts promptly and determine initial severity based on established procedures.

- Maintain accurate shift handover and operational records.

First-Level Incident Troubleshooting :

- Perform basic network connectivity checks including ping, DNS resolution and endpoint connectivity.

- Check OpenShift/Kubernetes resource status including nodes, pods, deployments and services.

- Review basic application and platform logs to identify common failure conditions.

- Check infrastructure and service health using approved dashboards and operational tools.

- Collect relevant diagnostic information before escalation.

Incident Recovery :

- Execute documented recovery procedures and operational runbooks.

- Restart or redeploy affected workloads using approved GitOps processes.

- Verify service recovery through dashboards, health checks and application endpoints.

- Escalate when recovery procedures are unsuccessful or when an incident falls outside the approved operating scope.

Incident Coordination & Escalation :

- Create and maintain incident tickets with accurate timestamps, symptoms, actions and observations.

- Engage the appropriate application, platform, infrastructure or network teams based on established escalation procedures.

- Provide clear status updates during active incidents.

- Support incident bridges by providing operational information and executing actions requested by L2/L3 engineers.

- Ensure effective handover of unresolved incidents between shifts.

Operational Procedures :

- Follow established Standard Operating Procedures (SOPs), runbooks and change-management processes.

- Document newly encountered symptoms and successful troubleshooting steps.

- Highlight recurring alerts or operational problems to senior platform engineers.

- Participate in operational drills and recovery exercises.

Scope of Authority :

Engineer may independently :

- Acknowledge and investigate alerts.

- Perform approved diagnostic commands and health checks.

- Execute documented L1 recovery procedures.

- Trigger approved GitOps redeployments.

- Open incidents and engage predefined support teams.

- Escalate incidents based on severity and runbook criteria.

Engineer must escalate :

- Changes requiring manual modification of production infrastructure.

- Unauthorised configuration or source-code changes.

- Security incidents or suspected compromises.

- Infrastructure failures requiring RE intervention.

- Incidents where documented recovery procedures fail.

- Major incidents requiring business or management decisions.

Qualifications & Experience :

- Bachelor's degree in Computer Science, Information Technology, Engineering or related discipline.

- 2 years of IT operations, infrastructure, cloud or application support experience.

- Basic Linux command-line knowledge.

- Basic networking knowledge including IP addressing, DNS, ports and connectivity troubleshooting.

- Basic understanding of containers and Kubernetes concepts.

- Ability to follow technical procedures accurately.

- Good written and verbal communication skills.

- Willingness and ability to work in a 24x7 shift environment.

Good to Have :

- Exposure to Kubernetes or Red Hat OpenShift.

- Exposure to Git and GitOps concepts.

- Familiarity with monitoring tools such as Grafana and Prometheus.

- Basic scripting experience with Bash or Python.

- Familiarity with incident-management or ITIL processes.

- Exposure to cloud or data-centre infrastructure.

- Interest in AI/ML infrastructure and GPU-based platforms.

Key Competencies :

- Systematic troubleshooting.

- Attention to detail.

- Ability to remain structured during incidents.

- Clear communication and escalation.

- Discipline in following operational procedures.

- Willingness to learn.

- Teamwork across geographically distributed support teams.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...