HamburgerMenu
hirist

Senior Site Reliability Engineer - Docker/Kubernetes

BharatHire.Com
8 - 13 Years
rupee25-35 LPA
Multiple Locations

Posted on: 29/08/2026

Job Description

Role Overview :


As a Senior Site Reliability Engineer, you will serve as a critical bridge between software development and IT operations, ensuring our platforms remain resilient, scalable, and highly available. You will work closely with cross-functional engineering teams, product managers, and infrastructure architects to design and implement robust systems that support our global user base. Your daily focus will involve automating manual processes, optimizing cloud infrastructure, and leading incident response efforts to minimize downtime. By championing reliability best practices, you will directly influence the performance of our services, ensuring a seamless experience for our customers and driving operational excellence across the organization.


Key Responsibilities :


- Architect and maintain highly available, fault-tolerant infrastructure on AWS and Azure to ensure consistent service delivery for our global clients.

- Automate infrastructure provisioning and configuration management using Terraform, Ansible, and CloudFormation to reduce deployment cycles and human error.



- Orchestrate containerized applications using Kubernetes and Docker to improve resource utilization and application scalability.



- Lead post-incident reviews and root cause analysis to implement long-term fixes that prevent recurring system failures.



- Mentor junior engineers and foster a culture of reliability by establishing clear SLOs, SLIs, and error budgets across development teams.



- Optimize cloud resource consumption and performance monitoring to balance cost-efficiency with high-performance requirements.


Required Skillset :


- Demonstrated expertise in managing large-scale production environments on AWS or Azure, with a deep understanding of cloud-native architecture.

- Advanced proficiency in container orchestration using Kubernetes and containerization via Docker, with a proven ability to troubleshoot complex cluster issues.



- Strong command over Infrastructure as Code (IaC) tools including Terraform, Ansible, and CloudFormation to manage complex, multi-region environments.



- Exceptional problem-solving skills and the ability to communicate technical complexities to non-technical stakeholders during high-pressure incidents.



- Proven experience in driving collaborative efforts across distributed teams, maintaining a proactive approach to system health and security.



- A Bachelor's or Master's degree in Computer Science or a related field, complemented by 8 to 13 years of hands-on experience in SRE or DevOps roles.



- Ability to thrive in a hybrid work environment, demonstrating self-motivation and the capacity to lead technical initiatives independently across our Gurgaon, Bangalore, Mumbai, or Chennai locations.


info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Posted by

Apoorva

TA at BharatHire.Com

Last Active: NA as recruiter has posted this job through third party tool.

Job Views:  
18
Applications:  3
Recruiter Actions:  0

Posted in

DevOps / SRE

Functional Area

Site Reliability Engineering

Job Code

1667071

Loading chat...