Posted on: 16/09/2026
Job Description :
- Own production platform reliability end-to-end.
- Kubernetes native multi-region platform.
- IC role.
- SLOs, observability & automation.
- On-call & incident program.
- Reliability of our EKS-based deployment platform - GitOps delivery.
- Hands-on engineering.
Required Skillset :
- Demonstrated expertise in managing large-scale AWS cloud infrastructure, with a deep understanding of EKS, networking, and security best practices.
- Proven ability to author and maintain complex Terraform or OpenTofu modules, ensuring infrastructure is scalable, modular, and version-controlled.
- Strong proficiency in Kubernetes orchestration, including troubleshooting complex cluster issues and optimizing resource utilization for high-traffic applications.
- Advanced experience in observability and incident management, specifically using Datadog to monitor distributed systems and maintain strict Service Level Objectives (SLOs).
- Exceptional communication skills, with the ability to articulate technical risks and architectural decisions to both engineering peers and non-technical stakeholders.
- A proactive, problem-solving mindset with the ability to thrive in a hybrid work environment in Gurugram, balancing independent deep work with collaborative team initiatives.
Total Experience : 5 - 8 years
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1671852