Posted on: 07/09/2026
Excellent MNC Opportunity - Immediate Joiner only
We're looking for a Senior Site Reliability Engineer to take ownership of reliability, performance, and operational excellence across our cloud infrastructure and Kubernetes platforms. You'll lead the response to complex, high-severity incidents, raise the bar on automation and observability, and mentor other engineers as the team scales.
What You'll Do :
- Lead troubleshooting and resolution of complex, high-severity incidents across AWS infrastructure and Amazon EKS clusters, often serving as incident commander for major incidents.
- Design and build Python-based automation frameworks and tooling that eliminate manual toil and improve reliability at scale, not just one-off scripts.
- Architect and maintain Terraform-based infrastructure-as-code, establishing reusable, secure, and scalable patterns across AWS environments.
- Drive observability strategy - design Grafana dashboards and alerting frameworks that surface the right signals to the right teams at the right time.
- Use SQL and CloudWatch Logs Insights to perform deep root-cause analysis on complex, cross-service incidents.
- Own end-to-end ITSM processes - Incident Management and Problem Management - including root cause analysis, post-incident reviews, and long-term remediation plans.
- Mentor junior and mid-level SREs, reviewing their troubleshooting approach, automation, and incident handling.
- Partner with engineering, product, and leadership teams to communicate incident impact, risk, and remediation plans clearly and confidently.
- Design and build self-healing automation and runbooks that detect known failure patterns and trigger remediation automatically, reducing manual intervention and recovery time for recurring incidents.
- Implement and maintain monitoring across multiple regions to ensure consistent visibility into system health, latency, and failover readiness across all deployment zones.
- Proactively identify potential failure points and performance bottlenecks before they impact production and reduce operational workload by automating recurring manual tasks.
- Participate in and provide senior-level escalation support for on-call rotations.
What We're Looking For :
- 8 - 10 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles, with a track record of owning reliability for production-critical systems.
- Deep, hands-on expertise with core and advanced AWS services (EC2, VPC, IAM, S3, RDS, CloudWatch, networking, etc.).
- Proven expertise troubleshooting complex Amazon EKS issues - cluster-level failures, networking, autoscaling, performance bottlenecks, and upgrade-related issues.
- Strong proficiency in Python for building automation frameworks, internal tooling, and operational systems.
- Extensive experience designing and maintaining Terraform modules and infrastructure patterns at scale.
- Strong command of ITSM frameworks, with hands-on ownership of Incident and Problem Management for high-severity issues.
- Advanced skills querying and analyzing data via SQL and AWS CloudWatch Logs Insights to drive root-cause analysis.
- Proven experience designing Grafana dashboards and alerting strategies that scale across multiple teams and services.
- Exceptional verbal and written communication skills - able to clearly articulate technical issues, risk, and remediation plans to engineering leadership and non-technical stakeholders alike.
- Experience mentoring or leading other engineers and contributing to team-level reliability strategy.
Mandatory Certifications :
- AWS Certification - required (e.g., AWS Certified Solutions Architect - Professional, AWS Certified DevOps Engineer - Professional, or equivalent).
- Certified Kubernetes Administrator (CKA) or equivalent EKS/Kubernetes certification - required.
Soft Skills :
- Calm, decisive leadership during high-pressure, high-severity incidents.
- A strong ownership mindset - drives issues to true resolution and follows through on long-term remediation.
- Natural mentor who raises the technical bar for the team.
- Collaborative cross-functional partner who works effectively with Dev, Infra, Product, and leadership.
Education :
- Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent extensive practical experience.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1669130