HamburgerMenu
hirist

Senior Site Reliability Engineer - Cloud Infrastructure

Ai Adept Consulting
8 - 13 Years
Bangalore

Posted on: 07/09/2026

Job Description

Excellent MNC Opportunity - Immediate Joiner only

We're looking for a Senior Site Reliability Engineer to take ownership of reliability, performance, and operational excellence across our cloud infrastructure and Kubernetes platforms. You'll lead the response to complex, high-severity incidents, raise the bar on automation and observability, and mentor other engineers as the team scales.

What You'll Do :

- Lead troubleshooting and resolution of complex, high-severity incidents across AWS infrastructure and Amazon EKS clusters, often serving as incident commander for major incidents.

- Design and build Python-based automation frameworks and tooling that eliminate manual toil and improve reliability at scale, not just one-off scripts.

- Architect and maintain Terraform-based infrastructure-as-code, establishing reusable, secure, and scalable patterns across AWS environments.

- Drive observability strategy - design Grafana dashboards and alerting frameworks that surface the right signals to the right teams at the right time.

- Use SQL and CloudWatch Logs Insights to perform deep root-cause analysis on complex, cross-service incidents.

- Own end-to-end ITSM processes - Incident Management and Problem Management - including root cause analysis, post-incident reviews, and long-term remediation plans.

- Mentor junior and mid-level SREs, reviewing their troubleshooting approach, automation, and incident handling.

- Partner with engineering, product, and leadership teams to communicate incident impact, risk, and remediation plans clearly and confidently.

- Design and build self-healing automation and runbooks that detect known failure patterns and trigger remediation automatically, reducing manual intervention and recovery time for recurring incidents.

- Implement and maintain monitoring across multiple regions to ensure consistent visibility into system health, latency, and failover readiness across all deployment zones.

- Proactively identify potential failure points and performance bottlenecks before they impact production and reduce operational workload by automating recurring manual tasks.

- Participate in and provide senior-level escalation support for on-call rotations.

What We're Looking For :

- 8 - 10 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles, with a track record of owning reliability for production-critical systems.

- Deep, hands-on expertise with core and advanced AWS services (EC2, VPC, IAM, S3, RDS, CloudWatch, networking, etc.).

- Proven expertise troubleshooting complex Amazon EKS issues - cluster-level failures, networking, autoscaling, performance bottlenecks, and upgrade-related issues.

- Strong proficiency in Python for building automation frameworks, internal tooling, and operational systems.

- Extensive experience designing and maintaining Terraform modules and infrastructure patterns at scale.

- Strong command of ITSM frameworks, with hands-on ownership of Incident and Problem Management for high-severity issues.

- Advanced skills querying and analyzing data via SQL and AWS CloudWatch Logs Insights to drive root-cause analysis.

- Proven experience designing Grafana dashboards and alerting strategies that scale across multiple teams and services.

- Exceptional verbal and written communication skills - able to clearly articulate technical issues, risk, and remediation plans to engineering leadership and non-technical stakeholders alike.

- Experience mentoring or leading other engineers and contributing to team-level reliability strategy.

Mandatory Certifications :

- AWS Certification - required (e.g., AWS Certified Solutions Architect - Professional, AWS Certified DevOps Engineer - Professional, or equivalent).

- Certified Kubernetes Administrator (CKA) or equivalent EKS/Kubernetes certification - required.

Soft Skills :

- Calm, decisive leadership during high-pressure, high-severity incidents.

- A strong ownership mindset - drives issues to true resolution and follows through on long-term remediation.

- Natural mentor who raises the technical bar for the team.

- Collaborative cross-functional partner who works effectively with Dev, Infra, Product, and leadership.

Education :

- Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent extensive practical experience.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...