HamburgerMenu
hirist

Job Description

Role Summary :

We are seeking an experienced Site Reliability Engineer (SRE) with strong hands-on expertise in cloud infrastructure, platform reliability, automation, observability, and production support.

This role focuses on improving service reliability through SLIs/SLOs and error budgets, reducing operational toil through automation, and partnering with global engineering teams to ensure resilient, scalable, and secure platforms.

Reliability Engineering :

- Define, measure, and report SLIs, SLOs, and error budgets for critical services.

- Drive service reliability improvements and systematically reduce operational toil through automation.

- Own capacity planning, performance tuning, and scalability initiatives.

- Lead blameless postmortems, root cause analyses, and corrective action tracking.

Platform Reliability & Automation :

- Design and maintain CI/CD pipelines using Azure DevOps and Jenkins.

- Operate and manage Azure cloud infrastructure and Kubernetes platforms.

- Deploy and support containerized applications using Docker, Kubernetes, and Helm.

- Automate infrastructure provisioning and configuration using Terraform and Ansible.

- Manage artifacts and repositories using JFrog Artifactory.

Observability & Production Support :

- Implement monitoring, logging, and observability using Prometheus, Grafana, Loki, and OpenTelemetry.

- Provide L2/L3 production support, incident management, troubleshooting, and RCA.

- Participate in on-call rotations supporting critical production services.

- Support PostgreSQL, Redis, and RabbitMQ environments, including high availability, backups, replication, and performance tuning.

Collaboration :

- Partner with Development, QA, Product, Operations, and global engineering teams.

- Ensure platform availability, scalability, security, and performance.

Mandatory Skills & Experience :

- 8 - 12 years of experience in Site Reliability Engineering.

- Strong hands-on experience with : Azure, Kubernetes, Docker, Jenkins, Terraform, Ansible.

- Proficiency in Python, Go, or Bash.

- Strong understanding of Networking & Security.

- Experience in L2/L3 production support, incident management, and root cause analysis.

- Experience working with global teams across multiple time zones.

- Strong communication, stakeholder management, and ownership mindset.

- Willingness to work from the Bangalore ITPL office, 5 days a week.

Good to Have :

- Experience with AI/GenAI concepts and AIOps practices.

- Exposure to chaos engineering and resilience testing.

- Azure (AZ-104/AZ-400), CKA, or CKAD certifications.

- Bachelor's degree in computer science or a related field.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...