HamburgerMenu
hirist

Job Description

Role Overview :

As the Director of Site Reliability Engineering, you will spearhead the vision, strategy, and operational excellence of our global IT infrastructure. You will lead high-performing engineering teams to ensure our platforms remain resilient, scalable, and highly available in a fast-paced, cloud-native environment.


Working closely with executive leadership, product managers, and cross-functional engineering heads, you will bridge the gap between development and operations to drive architectural improvements. Your leadership will directly influence the reliability of our services, ensuring a seamless experience for our customers while optimizing infrastructure costs and performance to support the company's long-term business objectives. This role is based in Hyderabad and requires a seasoned leader capable of navigating complex technical ecosystems.

Key Responsibilities :

- Define and execute the long-term SRE strategy to enhance system reliability and performance, ensuring that our infrastructure meets the rigorous demands of our growing user base.

- Lead and mentor large, distributed engineering teams, fostering a culture of continuous improvement, automation, and operational excellence across the organization.

- Oversee the design and maintenance of robust AWS-based cloud architectures, ensuring that our infrastructure is secure, cost-effective, and capable of handling massive scale.

- Drive the adoption of Kubernetes and container orchestration best practices to streamline deployment cycles and improve the efficiency of our microservices architecture.

- Optimize CI/CD pipelines to accelerate release velocity while maintaining strict quality gates, ensuring that product teams can ship features rapidly without compromising system stability.

- Manage complex incident response protocols and post-mortem processes, transforming technical failures into actionable insights that prevent future outages and improve system resilience.

Required Skillset :

- Demonstrated expertise in architecting and managing large-scale, mission-critical IT infrastructure within AWS environments, with a deep understanding of cloud-native design patterns.

- Proven ability to lead and scale high-performing engineering organizations, with a focus on talent development, performance management, and building inclusive, high-output teams.

- Extensive experience in implementing and managing Kubernetes clusters at scale, including deep knowledge of service meshes, observability, and infrastructure-as-code principles.

- Exceptional communication and stakeholder management skills, with the ability to translate complex technical challenges into strategic business outcomes for executive leadership.

- Strong background in CI/CD automation and DevOps methodologies, with a track record of reducing manual toil and improving deployment frequency through advanced tooling.

- A minimum of 18 - 25 years of experience in infrastructure engineering and reliability roles, ideally within high-growth technology companies or large-scale enterprise environments.

- Ability to thrive in a hybrid work environment in Hyderabad, demonstrating flexibility and a proactive approach to leading teams across different time zones and locations.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...