Posted on: 11/05/2026
Position : Site Reliability Engineer
Location : Bangalore .
NP : immediate joiner
Summary :
A Site Reliability Engineer (SRE) applies software engineering practices to operations to ensure services are reliable, scalable, and efficient. The SRE partner with product and platform teams to define SLIs/SLOs, automate operations, runbook and incident engineering, and drive long term reliability improvements.
Core Responsibilities :
- Service Reliability : Define SLIs/SLOs and monitor error budgets; drive actions when SLOs are at risk.
- Incident Management : Lead on call rotations, perform incident response, run postmortems, and implement corrective actions.
- Automation & Tooling : Build automation for deployment, remediation, scaling, and runbook tasks to reduce manual toil.
- Observability : Design and maintain metrics, logs, and distributed tracing to support rapid diagnosis and capacity planning.
- Performance & Capacity : Run load tests, capacity planning, and tuning to meet performance targets.
- Resilience Engineering : Implement canaries, chaos testing, circuit breakers, and progressive rollouts.
- Platform Improvement : Collaborate with dev teams to productionize features, harden services, and reduce operational risk.
- Knowledge Sharing : Produce runbooks, run regular reliability reviews, and coach teams on best practices.
Required Skills & Experience :
- Core : Strong programming/scripting (Python, Go, or equivalent), systems fundamentals (Linux), networking, and debugging skills.
- Observability : Experience with metrics platforms, logging, and tracing (Prometheus, Grafana, ELK/Opentelemetry or equivalent).
- Cloud & Infra : Familiarity with cloud platforms (AWS/Azure/GCP), containers, orchestration (Kubernetes), and IaC (Terraform/ARM).
- Operational Practice : Proven incident leadership, postmortem facilitation, and on call experience.
- Automation : CI/CD pipelines, deployment automation, and infrastructure automation experience.
- Soft skills : Clear communicator, calm under pressure, collaborative, and able to drive cross team change.
Preferred Qualifications :
- Bachelor's in computer science, Engineering, or equivalent experience.
- Experience defining SLIs/SLOs and using error budget processes.
- Prior SRE, production operations, or systems engineering role.
- Certifications or courses in cloud, Kubernetes, or reliability engineering (optional).
Success Metrics (KPIs) :
- Achieve and sustain defined SLO targets for owned services.
- Mean Time To Detect (MTTD) and Mean Time To Recover (MTTR) improvements quarter over quarter.
- Reduction in manual toil hours via automation.
- Number and quality of postmortem action items closed within SLA.
- Error budget burn rate and governance adherence.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1634726