Posted on: 22/05/2026
Job Title : Site Reliability Engineer (SRE) / Production Engineering Lead
Location : Hyderabad
Work Mode : Work From Office (5 days)
Experience : 12- 15 Years
Role Overview :
We are looking for an experienced Site Reliability Engineer (SRE) / Production Engineering Lead to manage and enhance the reliability, scalability, and performance of mission-critical systems. The role requires strong expertise in real-time payment systems (UPI/BFSI), production support, and platform engineering in high-availability environments.
Key Responsibilities :
- Ensure high availability, reliability, and performance of production systems
- Manage and lead P0/P1 incident bridges, ensuring quick resolution and minimal downtime
- Monitor system health using Prometheus, Grafana, and other observability tools
- Work closely with engineering teams to improve system resilience and scalability
- Perform root cause analysis (RCA) and implement preventive measures
- Support and optimize real-time payment systems (UPI/BFSI)
- Manage infrastructure across on-premises, data center, and hybrid environments
- Handle system components like Kafka, Redis, Nginx, and Tomcat for high-performance systems
- Drive automation initiatives for monitoring, deployment, and incident handling
Must-Have Requirements (Mandatory) :
- 12- 15 years of experience in SRE / Production Engineering / Platform roles
- Strong experience in UPI, Payments, or BFSI real-time systems (mandatory)
- Hands-on expertise with Prometheus and Grafana
- Proven experience handling critical production incidents (P0/P1 bridges)
- Experience working in on-premises / data center / hybrid environments
- Hands-on experience with Kafka, Redis, Nginx, Tomcat
- Strong expertise in Linux systems
Preferred Skills (Good to Have) :
- Experience with VictoriaMetrics, ElasticSearch, Jaeger
- Exposure to VM-based and Docker environments (non-Kubernetes)
- Experience in network debugging and performance tuning
- Automation experience using Shell scripting or Ansible
Key Competencies :
- Strong problem-solving and analytical mindset
- Ability to work in high-pressure, real-time environments
- Excellent communication and stakeholder management skills
- Proactive approach toward incident prevention and system optimization
Why Join :
- Opportunity to work on mission-critical, high-scale payment systems
- Exposure to real-time, high-impact production environments
- Collaborate with highly skilled teams in a core engineering setup
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1638192