HamburgerMenu
hirist

Site Reliability Engineering Lead

Careerist Management Consultants Pvt Ltd
9 - 15 Years
Multiple Locations

Posted on: 26/06/2026

Job Description

We are seeking an experienced and highly motivated Site Reliability Engineering (SRE) Lead to drive reliability, scalability, performance, and operational excellence across our cloud-native platforms and mission-critical applications.

The ideal candidate will possess strong expertise in Kubernetes, AWS, Infrastructure as Code, observability, automation, and production operations.

This role requires a hands-on technical leader who can build resilient systems, automate operational workflows, and mentor engineering teams in adopting SRE best practices.

Key Responsibilities :

- Design, implement, and maintain highly available, fault-tolerant, and scalable production platforms.

- Define and drive Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.

- Improve platform reliability, system uptime, and application performance through proactive engineering practices.

- Lead initiatives focused on reducing production incidents and minimizing downtime.

- Establish operational excellence standards across engineering teams.

- Architect and manage cloud-native infrastructure on AWS.

- Design and optimize Kubernetes-based container platforms for scalability, security, and performance.

- Implement infrastructure automation using Terraform and Infrastructure-as-Code (IaC) principles.

- Ensure high availability, disaster recovery, backup, and business continuity strategies are in place.

- Build comprehensive observability frameworks using Prometheus, Grafana, and related monitoring tools.

- Develop dashboards, alerting mechanisms, and health-check systems to improve operational visibility.

- Implement distributed tracing, log aggregation, and performance monitoring solutions.

- Analyze system metrics and proactively identify bottlenecks and reliability risks.

- Automate infrastructure provisioning, deployment, monitoring, and operational workflows.

- Develop scripts and automation tools using Python and shell scripting.

- Enhance CI/CD pipelines to support reliable, secure, and efficient application deployments.

- Reduce manual intervention through self-healing systems and automated remediation processes.

- Lead production incident response, troubleshooting, and root cause analysis (RCA).

- Establish incident management processes, runbooks, and operational playbooks.

- Conduct postmortems and implement preventive measures to avoid recurring issues.

- Drive continuous improvement initiatives based on production learnings.

- Mentor SREs, DevOps Engineers, and platform teams on reliability engineering best practices.

- Collaborate with Software Engineering, Architecture, Security, and Product teams.

- Lead architecture reviews with a focus on scalability, resiliency, and operational readiness.

- Promote a culture of automation, ownership, accountability, and continuous improvement.

Required Skills & Experience :

- Strong experience with Kubernetes administration, deployment, and troubleshooting.

- Expertise in Prometheus, Grafana, and observability platforms.

- Hands-on experience with Terraform for Infrastructure as Code.

- Strong Linux system administration and performance tuning expertise.

- Proficiency in Python scripting and automation.

- Extensive experience with AWS Cloud Services, including :

1. EC2

2. EKS

3. VPC

4. IAM

5. RDS

6. S3

7. CloudWatch

8. Route 53

9. Load Balancers

- Experience building and managing CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, or similar tools.

- Knowledge of containerization technologies such as Docker.

Reliability & Operations :

- Experience implementing SRE principles including SLOs, SLIs, and Error Budgets.

- Strong understanding of system reliability, scalability, and performance engineering.

- Expertise in incident management and production support processes.

- Experience leading SRE, DevOps, or Platform Engineering teams.

- Strong stakeholder management and cross-functional collaboration skills.

- Ability to drive technical decisions and reliability initiatives across multiple teams.

- Excellent communication, mentoring, and leadership capabilities.

- Experience managing large-scale distributed systems and microservices architectures.

- Knowledge of service mesh technologies such as Istio or Linkerd.

- Experience with Chaos Engineering and resilience testing.

- Exposure to security best practices for cloud-native environments.

- Understanding of FinOps and cloud cost optimization strategies.

- Experience with multi-cloud or hybrid-cloud environments.

- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.

- Improvement in platform uptime and service availability.

- Reduction in Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).

- Increased deployment frequency with reduced deployment failures.

- Enhanced observability coverage and operational visibility.

- Reduction in recurring incidents through automation and proactive reliability initiatives.

- Improved infrastructure scalability and operational efficiency.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...