Posted on: 10/09/2026
Job Description :
We are looking for a Lead Production Operations Engineer to lead the reliability, availability, security, and performance of critical production applications and cloud infrastructure. The role requires strong hands-on experience in GCP production environments, Kubernetes, application support, monitoring, incident management, and SRE practices, preferably within the BFSI domain.
- Lead production operations and support for critical business applications and cloud environments on GCP.
- Ensure high availability, reliability, scalability, and performance of production systems.
- Provide hands-on support for Kubernetes-based application environments, including deployment, troubleshooting, configuration, and operational management.
- Manage application incidents, service requests, problems, and major incidents, ensuring timely resolution and effective RCA.
- Monitor production environments using Dynatrace and other observability tools to identify performance, availability, and capacity issues.
- Drive proactive monitoring, alerting, automation, and operational improvements to reduce incidents and manual intervention.
- Troubleshoot application, infrastructure, Linux, networking, and connectivity issues across production environments.
- Implement and maintain DevOps and SRE practices to improve system reliability and operational efficiency.
- Ensure production environments follow appropriate security, access-control, and operational standards.
- Develop and maintain runbooks, operational procedures, knowledge documentation, and recovery processes.
- Coordinate with Development, Cloud, Infrastructure, Security, and other technology teams for production issues and service improvements.
- Support capacity planning, performance optimization, disaster recovery, and production readiness activities.
- Drive continuous improvement of IT service management processes in line with ITIL practices.
- Lead and mentor production support teams and manage operational priorities across critical applications.
Required Skills :
- 11 - 14 years of experience in Production Operations, Application Support, SRE, DevOps, or Cloud Operations.
- 7+ years of hands-on experience supporting production environments on GCP.
- 3 - 5 years of hands-on Kubernetes experience in production.
- Strong experience in application production support, incident management, problem management, and RCA.
- Strong Linux administration and troubleshooting skills.
- Experience with Dynatrace or similar monitoring and observability platforms.
- Strong understanding of networking, security, cloud infrastructure, and production troubleshooting.
- Experience with DevOps and SRE practices, automation, and operational reliability.
- Good understanding of ITIL-based service management processes.
- Experience working with highly available, business-critical production applications.
- BFSI domain experience is preferred.
- Strong leadership, stakeholder management, communication, and problem-solving skills.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1670283