Posted on: 05/06/2026
Role Overview :
As a seasoned Site Reliability Engineer, you will be instrumental in ensuring the unwavering availability, performance, and scalability of our critical production systems. This role demands a proactive approach to operational excellence, working closely with our development, product, and infrastructure teams to embed reliability principles throughout the software development lifecycle.
You will champion a culture of automation and continuous improvement, directly impacting the stability of our services, enhancing user trust, and safeguarding business continuity for our diverse customer base.
Key Responsibilities :
- Lead the architectural design and implementation of highly resilient and scalable infrastructure solutions, focusing on Service Reliability and performance optimization for critical applications.
- Drive the development and enforcement of robust Disaster Recovery and IT Disaster Recovery plans, conducting regular drills and ensuring swift, effective recovery strategies to minimize downtime.
- Define, monitor, and optimize Service Level Indicators (SLIs) and Service Level Objectives (SLOs), establishing `Error Budgets` to balance innovation with operational stability and meet stringent `SLA` commitments.
- Spearhead Incident Management processes, leading post-incident reviews to identify root causes, implement corrective actions, and prevent recurrence, fostering a learning culture.
- Develop and implement advanced Troubleshooting methodologies and tools to quickly diagnose and resolve complex system issues across distributed environments.
- Automate operational tasks, infrastructure provisioning, and deployment pipelines using modern SRE practices, reducing manual toil and improving system efficiency.
- Provide expert guidance and mentorship in System Administration best practices, cloud infrastructure, and observability tools to cross-functional teams, elevating overall technical capabilities.
- Collaborate with engineering teams to design and build systems that are inherently observable, scalable, and maintainable from inception.
- Evaluate and integrate new technologies and tools to enhance our reliability posture, security, and operational efficiency.
Required Skillset :
- Demonstrated expertise in designing, building, and operating large-scale, highly available distributed systems, with a deep understanding of Site Reliability principles.
- Proven ability to lead and execute complex Disaster Recovery and IT Disaster Recovery initiatives, including strategy formulation, testing, and execution.
- Strong analytical skills to define, track, and interpret Service Level Indicators (SLIs), manage Error Budgets, and ensure SLA adherence.
- Exceptional Troubleshooting and problem-solving capabilities for intricate production issues across diverse technology stacks and cloud environments.
- Extensive experience in System Administration for Linux-based systems, including networking, security, and performance tuning.
- Proficient in scripting and automation (e.g., Python, Go, Shell) and experience with infrastructure-as-code tools (e.g., Terraform, Ansible).
- Strong leadership and Incident Management skills, with the ability to remain calm and decisive under pressure during critical outages.
- Excellent communication and interpersonal skills, capable of articulating complex technical concepts to both technical and non-technical stakeholders.
- Bachelor's or Master's degree in Computer Science, Engineering, or a related technical field from a reputable institution.
- Adaptability to work in a dynamic, fast-paced environment based out of our Navi Mumbai office, collaborating effectively with global teams.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1642043