HamburgerMenu
hirist

Lead Production Operations Engineer - Site Reliability

Quotient Consultancy
11 - 14 Years
Guwahati

Posted on: 09/09/2026

Job Description

Job Summary :

We are looking for an experienced Lead Production Operations Engineer to ensure the reliability, scalability, security, and performance of production applications and cloud infrastructure.

The role will provide L3 production support, drive SRE and DevOps best practices, and collaborate with Engineering, Security, Incident Management, and Product teams.

Key Responsibilities :

- Own production operations, application support, incident resolution, and service reliability.

- Provide L3 technical support and troubleshoot complex production issues.

- Implement SRE, DevOps, monitoring, alerting, logging, and reliability best practices.

- Manage and support production workloads on GCP and Kubernetes.

- Monitor application and infrastructure performance using Dynatrace and other observability tools.

- Identify operational risks, perform root-cause analysis, and implement preventive and corrective actions.

- Manage cloud infrastructure components including networking, security, load balancing, autoscaling, databases, backup, and recovery.

- Collaborate with Development, Security, Product, Business, and Incident Management teams to ensure stable production environments.

- Drive continuous improvement, automation, operational efficiency, and technology adoption.

- Ensure adherence to ITIL, security, compliance, and operational standards.

Required Skills & Experience :

- 11 - 14 years of overall experience in Production/Application Operations or Technical Support.

- 7+ years of production support experience on Google Cloud Platform (GCP).

- 3 - 5 years of hands-on Kubernetes experience.

- Strong experience in Application Support, Cloud Infrastructure, Linux, Networking, and DevOps.

- Hands-on experience with GCP services, networking, security, load balancers, autoscaling, databases, backup and recovery.

- Strong understanding of Kubernetes, containerization, container orchestration, and container registries.

- Experience with logging, monitoring, alerting, and observability, preferably Dynatrace.

- Basic understanding of SRE principles and reliability engineering.

- Experience with ITIL-based production support and incident management.

- BFSI domain experience preferred.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...