Posted on: 13/07/2026
Role Overview :
We are seeking a seasoned Lead Site Reliability Engineer to spearhead our infrastructure reliability initiatives and drive the evolution of our production environments.
In this role, you will act as the technical anchor for our SRE practice, working closely with cross-functional engineering teams, product managers, and executive stakeholders to define and uphold rigorous service level objectives.
You will be responsible for building resilient, scalable systems that directly impact the end-user experience, ensuring that our platform remains performant and available under high-traffic conditions.
By fostering a culture of automation and proactive monitoring, you will play a pivotal role in minimizing technical debt and accelerating our deployment velocity in a fully remote environment.
Key Responsibilities :
- Architect and maintain highly available cloud infrastructure on AWS to ensure seamless service delivery for our global customer base.
- Lead the design and implementation of robust automation frameworks to eliminate manual toil and improve the efficiency of our deployment pipelines.
- Mentor a team of engineers by providing technical guidance on complex system troubleshooting and performance tuning to maintain high operational standards.
- Collaborate with development teams to integrate observability and monitoring tools, ensuring proactive identification and resolution of potential system bottlenecks.
- Establish and enforce best practices for incident management and post-mortem analysis to continuously improve system stability and team learning.
Required Skillset :
- Demonstrated expertise in managing large-scale distributed systems with 8 - 14 years of professional experience in SRE or DevOps environments.
- Advanced proficiency in Linux system administration and deep technical knowledge of AWS cloud services, including networking, security, and storage optimization.
- Proven ability to design and execute comprehensive automation testing strategies that ensure the reliability of infrastructure-as-code deployments.
- Strong interpersonal skills with the ability to communicate complex technical concepts to non-technical stakeholders and influence cross-team architectural decisions.
- Exceptional capacity for self-management and collaborative problem-solving within a distributed, remote-first team structure.
- A Bachelors or Masters degree in Computer Science or a related field, reflecting a strong foundation in engineering principles.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1653502