Posted on: 08/07/2026
Role : Lead Site Reliability Engineer (Agentic AI)
Location : Hyderabad, India | Full-Time | On-site / Hybrid
Experience : 69 Years
The Role :
Own reliability engineering for mission-critical healthcare AI infrastructure and define the future of our SRE practice.
You'll architect highly available, secure, and scalable systems supporting global AI workloads.
What You'll Do :
Reliability Engineering :
- Define and own :
1. SLOs
2. SLIs
3. Error Budgets
4. Incident Management
- Improve platform availability and operational excellence.
Infrastructure Architecture :
- Architect large-scale cloud infrastructure.
- Lead Kubernetes platform engineering.
- Drive disaster recovery and capacity planning.
Observability :
- Build and manage :
1. Prometheus
2. Grafana
3. OpenTelemetry
4. ELK
5. Distributed Tracing systems
AI Systems Reliability :
- Operate GPU infrastructure.
- Improve LLM serving reliability.
- Build self-healing systems.
Leadership :
- Mentor engineers.
- Lead incident management.
- Drive SRE best practices across the organization.
What We're Looking For :
- 6 to 9 years in SRE, DevOps, Platform Engineering, or MLOps.
- Deep expertise in :
1. AWS
2. GCP
3. Kubernetes
4. Terraform
5. Linux
6. Networking
- Strong programming skills :
1. Python
2. Go
3. Bash
- Proven track record operating large-scale production systems and AI infrastructure.
Why Join Us? :
- Build infrastructure powering real-world healthcare AI.
- Work with exceptional AI-native engineering teams.
- Solve complex distributed systems challenges with significant ownership and impact.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
DevOps / Cloud
Job Code
1652347