HamburgerMenu
hirist

Senior Site Reliability Engineer - Cloud Infrastructure

BharatHire.Com
8 - 13 Years
Multiple Locations

Posted on: 10/09/2026

Job Description

Role Overview :

We are looking for a Senior Site Reliability Engineer with strong experience in cloud-based production environments, Kubernetes, infrastructure automation, observability, incident management, and reliability engineering. The role will focus on maintaining highly available and scalable production systems, automating operational processes, and driving continuous improvements in reliability and performance.

Key Responsibilities :

- Design, deploy, and operate highly available and scalable cloud-based production environments.

- Manage and troubleshoot containerized workloads using Docker and Kubernetes.

- Implement Infrastructure as Code using Terraform, Ansible, CloudFormation, or similar tools.

- Build and maintain monitoring, logging, alerting, and observability solutions for production systems.

- Define and monitor reliability metrics such as SLIs, SLOs, SLAs, and error budgets.

- Lead production incident response, troubleshooting, root-cause analysis, post-incident reviews, and corrective actions.

- Automate repetitive operational activities, deployment processes, infrastructure provisioning, and incident-resolution workflows using scripting and automation.

- Manage and troubleshoot distributed systems and platforms including Kafka and relational/NoSQL databases.

- Support CI/CD pipelines and deployment processes with a focus on reliability, rollback, and operational safety.

- Identify performance, scalability, availability, and capacity issues and implement appropriate improvements.

- Develop and maintain runbooks, operational documentation, and reliability standards.

- Collaborate with Development, Platform, Security, and Infrastructure teams to improve production readiness and system resilience.

- Participate in on-call support and ensure timely resolution of critical production issues.

Required Skills :

- 8 - 10 years of experience in SRE, DevOps, Platform Engineering, Cloud Infrastructure, or related roles.

- Strong hands-on experience with Kubernetes and Docker in production environments.

- Strong experience with cloud platforms such as AWS, Azure, or GCP.

- Hands-on experience with Terraform, Ansible, CloudFormation, or equivalent IaC tools.

- Strong knowledge of monitoring, logging, alerting, and observability platforms.

- Experience with Kafka and production database environments.

- Strong scripting/automation skills using Python, Bash, Shell, or similar languages.

- Proven experience in incident management, RCA, problem management, and production troubleshooting.

- Understanding of SRE principles, SLI/SLO/SLA, error budgets, availability, scalability, and reliability engineering.

- Experience with CI/CD, deployment automation, and version control.

- Strong understanding of Linux, networking, security, and cloud-native environments.

- Strong analytical, troubleshooting, communication, and stakeholder-management skills.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...