HamburgerMenu
hirist

Job Description

Role : Lead Site Reliability Engineer (Agentic AI)

Location : Hyderabad, India | Full-Time | On-site / Hybrid

Experience : 69 Years

The Role :

Own reliability engineering for mission-critical healthcare AI infrastructure and define the future of our SRE practice.

You'll architect highly available, secure, and scalable systems supporting global AI workloads.

What You'll Do :

Reliability Engineering :

- Define and own :

1. SLOs

2. SLIs

3. Error Budgets

4. Incident Management

- Improve platform availability and operational excellence.

Infrastructure Architecture :

- Architect large-scale cloud infrastructure.

- Lead Kubernetes platform engineering.

- Drive disaster recovery and capacity planning.

Observability :

- Build and manage :

1. Prometheus

2. Grafana

3. OpenTelemetry

4. ELK

5. Distributed Tracing systems

AI Systems Reliability :

- Operate GPU infrastructure.

- Improve LLM serving reliability.

- Build self-healing systems.

Leadership :

- Mentor engineers.

- Lead incident management.

- Drive SRE best practices across the organization.

What We're Looking For :

- 6 to 9 years in SRE, DevOps, Platform Engineering, or MLOps.

- Deep expertise in :

1. AWS

2. GCP

3. Kubernetes

4. Terraform

5. Linux

6. Networking

- Strong programming skills :

1. Python

2. Go

3. Bash

- Proven track record operating large-scale production systems and AI infrastructure.

Why Join Us? :

- Build infrastructure powering real-world healthcare AI.

- Work with exceptional AI-native engineering teams.

- Solve complex distributed systems challenges with significant ownership and impact.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...