Posted on: 19/09/2026
Role Overview :
As the Director of Site Reliability Engineering, you will spearhead the vision, strategy, and operational excellence of our global IT infrastructure. You will lead high-performing engineering teams to ensure our platforms remain resilient, scalable, and highly available in a fast-paced, cloud-native environment.
Working closely with executive leadership, product managers, and cross-functional engineering heads, you will bridge the gap between development and operations to drive architectural improvements. Your leadership will directly influence the reliability of our services, ensuring a seamless experience for our customers while optimizing infrastructure costs and performance to support the company's long-term business objectives. This role is based in Hyderabad and requires a seasoned leader capable of navigating complex technical ecosystems.
Key Responsibilities :
- Define and execute the long-term SRE strategy to enhance system reliability and performance, ensuring that our infrastructure meets the rigorous demands of our growing user base.
- Lead and mentor large, distributed engineering teams, fostering a culture of continuous improvement, automation, and operational excellence across the organization.
- Oversee the design and maintenance of robust AWS-based cloud architectures, ensuring that our infrastructure is secure, cost-effective, and capable of handling massive scale.
- Drive the adoption of Kubernetes and container orchestration best practices to streamline deployment cycles and improve the efficiency of our microservices architecture.
- Optimize CI/CD pipelines to accelerate release velocity while maintaining strict quality gates, ensuring that product teams can ship features rapidly without compromising system stability.
- Manage complex incident response protocols and post-mortem processes, transforming technical failures into actionable insights that prevent future outages and improve system resilience.
Required Skillset :
- Demonstrated expertise in architecting and managing large-scale, mission-critical IT infrastructure within AWS environments, with a deep understanding of cloud-native design patterns.
- Proven ability to lead and scale high-performing engineering organizations, with a focus on talent development, performance management, and building inclusive, high-output teams.
- Extensive experience in implementing and managing Kubernetes clusters at scale, including deep knowledge of service meshes, observability, and infrastructure-as-code principles.
- Exceptional communication and stakeholder management skills, with the ability to translate complex technical challenges into strategic business outcomes for executive leadership.
- Strong background in CI/CD automation and DevOps methodologies, with a track record of reducing manual toil and improving deployment frequency through advanced tooling.
- A minimum of 18 - 25 years of experience in infrastructure engineering and reliability roles, ideally within high-growth technology companies or large-scale enterprise environments.
- Ability to thrive in a hybrid work environment in Hyderabad, demonstrating flexibility and a proactive approach to leading teams across different time zones and locations.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1672789