Posted on: 03/06/2026
Job Description :
We are looking for a Senior DevOps Engineer to drive the design, reliability, scalability, and operational excellence of our cloud infrastructure. This role demands strong ownership across infrastructure architecture, CI/CD, observability, automation, and large-scale production environments.
Key Responsibilities :
- Own the design, architecture, and reliability of cloud infrastructure across AWS, Azure, GCP, and Aliyun for global, multi-region deployments
- Lead and optimize the CI/CD ecosystem, including scaling and improving Jenkins-as-Code environments
- Drive the organizations Infrastructure as Code (IaC) journey by codifying cloud resources, alarms, and configurations with strong version control and rollback practices
- Collaborate with engineering teams to diagnose and resolve performance, scalability, and reliability challenges
- Define and implement best practices for monitoring, alerting, observability, and incident response
- Lead initiatives focused on cost optimization, security hardening, and capacity planning
- Mentor junior engineers and elevate DevOps maturity across teams
Required Skills & Experience :
- 5+ years of experience in DevOps, SRE, or Infrastructure Engineering
- Strong hands-on experience operating large-scale production environments with clear exposure to traffic, uptime, latency, or infrastructure scale
- Experience working in B2B SaaS product companies with multi-tenant architectures, multi-environment setups, or multi-client systems
- Strong expertise in AWS, including VPC, EKS, EC2, RDS, and networking
- Strong hands-on experience with Kubernetes (EKS) and designing high-availability, multi-region systems
- Strong experience with Terraform, Helm, GitOps, and Infrastructure as Code practices
- Strong scripting expertise in Python, Go, or Bash
- Experience building and managing scalable CI/CD pipelines using Jenkins, GitHub Actions, or similar tools
- Experience with zero or low-downtime deployment strategies
- Strong understanding of SRE principles, including SLOs, SLIs, and error budgets
- Hands-on experience with Prometheus, Grafana, Datadog, and modern observability practices
- Strong exposure to alerting, on-call operations, and incident management
Eligibility Criteria :
- B.Tech in Computer Science or related fields
- Experience from strong, scaled B2B SaaS product companies only
Preferred Profile :
- Strong ownership mindset with the ability to design, scale, and operate production infrastructure
- Experience mentoring engineers and driving engineering best practices
- Ability to work across teams to improve platform reliability, efficiency, and operational readiness
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
DevOps / Cloud
Job Code
1641167