HamburgerMenu
hirist

Lead Site Reliability Engineer - AWS/Kubernetes

Neemtree
6 - 10 Years
Mumbai

Posted on: 08/09/2026

Job Description

Job Overview:

We are seeking a highly motivated and strategic-minded Cloud Engineer to join our dynamic team. You will lead our reliability strategy, build self-healing architecture, eliminate operational toil through software engineering, and mentor a high-performing team of SRE and Devops engineers.

Responsibilities:

Architecture & Reliability Ownership:

- Design multi-AZ and multi-region AWS architectures for zero-downtime releases.

- Establish and enforce SLIs, SLOs, and error budgets alongside product managers.

Infrastructure as Code (IaC):

- Standardize scalable cloud infrastructure using Terraform or AWS CDK.

- Ensure state isolation, modularity, and automated deployment pipelines.

Observability & Monitoring:

- Drive signal-to-noise alerting and build observability pipelines using Prometheus/Grafana, APM, AWS CloudWatch, or OpenTelemetry.

Incident Leadership & Postmortems:

- Serve as Incident Commander for high-severity issues.

- Lead blameless post-mortems and convert system failures into concrete engineering work.

Automation & Platform Tooling:

- Treat operations as a software problem by writing custom tooling in Bash or Python to eliminate manual operational toil.

Developer Experience & CI/CD:

- Maintain robust CI/CD pipelines (GitHub Actions, Tekton ArgoCD) to empower product developers with safe self-service deployments.

FinOps & Security:

- Partner with finance and security to optimize AWS spending (Savings Plans, right-sizing) while ensuring IAM policies, KMS encryption, and VPC configurations meet compliance standards.

Mentorship & On-Call Management:

- Set sustainable on-call rotation policies to prevent burnout and mentor junior and mid-level SREs.

Qualifications:

- Experience: 6 - 10 years in DevOps, Platform Engineering, or Infrastructure.

- AWS Mastery: Deep hands-on experience across core AWS services (EKS, Lambda, RDS S3, Route 53, Transit Gateway).

- Kubernetes Expertise: Experience managing production EKS clusters, including workload scaling, ingress control, service meshes, and GitOps deployments (ArgoCD/Flux).

- Software Development: Strong coding skills in Bash or Python to build CLI tools, custom controllers, and API integrations.

- IaC Proficiency: Proven production usage of Terraform or AWS CDK.

- Observability: Track record of building metrics, logs, and trace pipelines from the ground up using modern observability stacks.

- Relevant certifications (e.g., AWS Certified DevOps Engineer, Certified Kubernetes Administrator) are a plus.

If you are passionate about AWS cloud technologies, databases, and automation, and want to be part of a dynamic team, apply now!

Experience Range : 6 - 10 years

Educational Qualifications : B. Tech/B. E, and M. Tech

Skills Required : AWS, EKS, Argocd, Flux cd, Bash, Python, Terraforms, Kubernetes

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...