Posted on: 26/06/2026
We are seeking an experienced and highly motivated Site Reliability Engineering (SRE) Lead to drive reliability, scalability, performance, and operational excellence across our cloud-native platforms and mission-critical applications.
The ideal candidate will possess strong expertise in Kubernetes, AWS, Infrastructure as Code, observability, automation, and production operations.
This role requires a hands-on technical leader who can build resilient systems, automate operational workflows, and mentor engineering teams in adopting SRE best practices.
Key Responsibilities :
- Design, implement, and maintain highly available, fault-tolerant, and scalable production platforms.
- Define and drive Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
- Improve platform reliability, system uptime, and application performance through proactive engineering practices.
- Lead initiatives focused on reducing production incidents and minimizing downtime.
- Establish operational excellence standards across engineering teams.
- Architect and manage cloud-native infrastructure on AWS.
- Design and optimize Kubernetes-based container platforms for scalability, security, and performance.
- Implement infrastructure automation using Terraform and Infrastructure-as-Code (IaC) principles.
- Ensure high availability, disaster recovery, backup, and business continuity strategies are in place.
- Build comprehensive observability frameworks using Prometheus, Grafana, and related monitoring tools.
- Develop dashboards, alerting mechanisms, and health-check systems to improve operational visibility.
- Implement distributed tracing, log aggregation, and performance monitoring solutions.
- Analyze system metrics and proactively identify bottlenecks and reliability risks.
- Automate infrastructure provisioning, deployment, monitoring, and operational workflows.
- Develop scripts and automation tools using Python and shell scripting.
- Enhance CI/CD pipelines to support reliable, secure, and efficient application deployments.
- Reduce manual intervention through self-healing systems and automated remediation processes.
- Lead production incident response, troubleshooting, and root cause analysis (RCA).
- Establish incident management processes, runbooks, and operational playbooks.
- Conduct postmortems and implement preventive measures to avoid recurring issues.
- Drive continuous improvement initiatives based on production learnings.
- Mentor SREs, DevOps Engineers, and platform teams on reliability engineering best practices.
- Collaborate with Software Engineering, Architecture, Security, and Product teams.
- Lead architecture reviews with a focus on scalability, resiliency, and operational readiness.
- Promote a culture of automation, ownership, accountability, and continuous improvement.
Required Skills & Experience :
- Strong experience with Kubernetes administration, deployment, and troubleshooting.
- Expertise in Prometheus, Grafana, and observability platforms.
- Hands-on experience with Terraform for Infrastructure as Code.
- Strong Linux system administration and performance tuning expertise.
- Proficiency in Python scripting and automation.
- Extensive experience with AWS Cloud Services, including :
1. EC2
2. EKS
3. VPC
4. IAM
5. RDS
6. S3
7. CloudWatch
8. Route 53
9. Load Balancers
- Experience building and managing CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, or similar tools.
- Knowledge of containerization technologies such as Docker.
Reliability & Operations :
- Experience implementing SRE principles including SLOs, SLIs, and Error Budgets.
- Strong understanding of system reliability, scalability, and performance engineering.
- Expertise in incident management and production support processes.
- Experience leading SRE, DevOps, or Platform Engineering teams.
- Strong stakeholder management and cross-functional collaboration skills.
- Ability to drive technical decisions and reliability initiatives across multiple teams.
- Excellent communication, mentoring, and leadership capabilities.
- Experience managing large-scale distributed systems and microservices architectures.
- Knowledge of service mesh technologies such as Istio or Linkerd.
- Experience with Chaos Engineering and resilience testing.
- Exposure to security best practices for cloud-native environments.
- Understanding of FinOps and cloud cost optimization strategies.
- Experience with multi-cloud or hybrid-cloud environments.
- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.
- Improvement in platform uptime and service availability.
- Reduction in Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).
- Increased deployment frequency with reduced deployment failures.
- Enhanced observability coverage and operational visibility.
- Reduction in recurring incidents through automation and proactive reliability initiatives.
- Improved infrastructure scalability and operational efficiency.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1648790