Posted on: 29/08/2026
Role Overview :
As a Senior Site Reliability Engineer, you will serve as a critical bridge between software development and IT operations, ensuring our platforms remain resilient, scalable, and highly available. You will work closely with cross-functional engineering teams, product managers, and infrastructure architects to design and implement robust systems that support our global user base. Your daily focus will involve automating manual processes, optimizing cloud infrastructure, and leading incident response efforts to minimize downtime. By championing reliability best practices, you will directly influence the performance of our services, ensuring a seamless experience for our customers and driving operational excellence across the organization.
Key Responsibilities :
- Automate infrastructure provisioning and configuration management using Terraform, Ansible, and CloudFormation to reduce deployment cycles and human error.
- Orchestrate containerized applications using Kubernetes and Docker to improve resource utilization and application scalability.
- Lead post-incident reviews and root cause analysis to implement long-term fixes that prevent recurring system failures.
- Mentor junior engineers and foster a culture of reliability by establishing clear SLOs, SLIs, and error budgets across development teams.
- Optimize cloud resource consumption and performance monitoring to balance cost-efficiency with high-performance requirements.
Required Skillset :
- Advanced proficiency in container orchestration using Kubernetes and containerization via Docker, with a proven ability to troubleshoot complex cluster issues.
- Strong command over Infrastructure as Code (IaC) tools including Terraform, Ansible, and CloudFormation to manage complex, multi-region environments.
- Exceptional problem-solving skills and the ability to communicate technical complexities to non-technical stakeholders during high-pressure incidents.
- Proven experience in driving collaborative efforts across distributed teams, maintaining a proactive approach to system health and security.
- A Bachelor's or Master's degree in Computer Science or a related field, complemented by 8 to 13 years of hands-on experience in SRE or DevOps roles.
- Ability to thrive in a hybrid work environment, demonstrating self-motivation and the capacity to lead technical initiatives independently across our Gurgaon, Bangalore, Mumbai, or Chennai locations.
Did you find something suspicious?
Posted by
Apoorva
Last Active: NA as recruiter has posted this job through third party tool.
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1667071