Posted on: 21/09/2026
Role Overview :
As a Senior DevOps/SRE Engineer based in Gurugram, you will be the backbone of our infrastructure reliability and scalability efforts. You will work closely with cross-functional engineering teams to design, deploy, and maintain high-availability systems that serve our global user base.
Your day-to-day will involve architecting robust cloud environments, automating deployment pipelines, and proactively monitoring system health to ensure seamless performance. By bridging the gap between development and operations, you will directly influence our ability to ship features faster while maintaining the rock-solid stability our customers expect, ultimately driving business growth through operational excellence.
Key Responsibilities :
- Architect and manage scalable AWS cloud infrastructure to ensure high availability and cost-efficiency for our production environments.
- Design and implement complex Kubernetes clusters to orchestrate containerized applications, improving deployment consistency across environments.
- Develop and maintain Infrastructure as Code (IaC) using Terraform to enable rapid, repeatable, and version-controlled environment provisioning.
- Build and optimize CI/CD pipelines to accelerate the software development lifecycle and reduce manual intervention during releases.
- Establish comprehensive monitoring and alerting frameworks using Prometheus and Grafana to provide real-time visibility into system performance and incident response.
- Perform deep-dive troubleshooting and performance tuning on Linux-based systems to resolve bottlenecks and optimize resource utilization for end-users.
Required Skillset :
- Demonstrated expertise in managing large-scale AWS environments and orchestrating containerized workloads using Kubernetes in a production setting.
- Proven ability to write clean, modular Terraform code to automate infrastructure provisioning and lifecycle management.
- Strong proficiency in Linux system administration, including shell scripting and kernel-level performance tuning to ensure system stability.
- Deep experience in configuring Prometheus and Grafana to create actionable dashboards and automated alerting systems that minimize downtime.
- Exceptional communication skills, with the ability to articulate complex technical challenges to stakeholders and mentor junior team members.
- A collaborative mindset, comfortable working in a hybrid office environment in Gurugram, with the adaptability to thrive in a fast-paced, high-growth startup culture.
- A Bachelor's degree in Computer Science, Engineering, or a related field, reflecting a strong foundation in distributed systems and software engineering principles.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1673021