Posted on: 09/09/2026
Job Summary :
We are looking for an experienced Lead Production Operations Engineer to ensure the reliability, scalability, security, and performance of production applications and cloud infrastructure.
The role will provide L3 production support, drive SRE and DevOps best practices, and collaborate with Engineering, Security, Incident Management, and Product teams.
Key Responsibilities :
- Own production operations, application support, incident resolution, and service reliability.
- Provide L3 technical support and troubleshoot complex production issues.
- Implement SRE, DevOps, monitoring, alerting, logging, and reliability best practices.
- Manage and support production workloads on GCP and Kubernetes.
- Monitor application and infrastructure performance using Dynatrace and other observability tools.
- Identify operational risks, perform root-cause analysis, and implement preventive and corrective actions.
- Manage cloud infrastructure components including networking, security, load balancing, autoscaling, databases, backup, and recovery.
- Collaborate with Development, Security, Product, Business, and Incident Management teams to ensure stable production environments.
- Drive continuous improvement, automation, operational efficiency, and technology adoption.
- Ensure adherence to ITIL, security, compliance, and operational standards.
Required Skills & Experience :
- 11 - 14 years of overall experience in Production/Application Operations or Technical Support.
- 7+ years of production support experience on Google Cloud Platform (GCP).
- 3 - 5 years of hands-on Kubernetes experience.
- Strong experience in Application Support, Cloud Infrastructure, Linux, Networking, and DevOps.
- Hands-on experience with GCP services, networking, security, load balancers, autoscaling, databases, backup and recovery.
- Strong understanding of Kubernetes, containerization, container orchestration, and container registries.
- Experience with logging, monitoring, alerting, and observability, preferably Dynatrace.
- Basic understanding of SRE principles and reliability engineering.
- Experience with ITIL-based production support and incident management.
- BFSI domain experience preferred.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1670057