Posted on: 15/05/2026
Roles & Responsibilities :
- Design, build, and evolve cloud-native infrastructure across AWS and Kubernetes environments to support global, multi-region deployments.
- Own the scalability, reliability, availability, and cost optimization of infrastructure across products and platforms.
- Build internal platforms, automation frameworks, and self-service developer tools that improve engineering productivity and operational efficiency.
- Architect and maintain secure, scalable, and high-performance CI/CD pipelines using GitHub Actions and Jenkins.
- Drive Infrastructure as Code (IaC) adoption using Terraform and improve deployment automation through Helm and GitOps practices.
- Implement zero or low-downtime deployment strategies and improve release engineering processes.
- Champion SRE practices including SLOs, SLIs, error budgets, incident response, and operational readiness.
- Build and enhance observability frameworks using tools such as Prometheus, Grafana, and Datadog.
- Lead troubleshooting and resolution of infrastructure, networking, scalability, and production reliability issues.
- Drive DevSecOps initiatives including policy-as-code, secrets management, vulnerability scanning, and infrastructure governance.
- Partner closely with engineering and platform teams on architecture decisions, deployment patterns, and production readiness.
- Mentor DevOps and SRE engineers and contribute to improving overall platform engineering maturity across teams.
Ideal Candidate Criteria :
- 9+ years of hands-on experience in DevOps, SRE, Infrastructure Engineering, or Platform Engineering roles.
- Experience working in Staff, Lead, or Senior DevOps/SRE roles with ownership across infrastructure and platform systems.
- Strong experience in B2B SaaS product companies with multi-tenant architectures or multiple production environments.
- Proven exposure to large-scale infrastructure environments with clear experience handling uptime, latency, scalability, and distributed systems challenges.
- Strong expertise in AWS services including VPC, EKS, EC2, RDS, networking, and cloud security.
- Hands-on experience managing Kubernetes (EKS) environments at scale.
- Strong understanding of high availability systems, disaster recovery, and multi-region architecture design.
- Deep expertise in Terraform for Infrastructure as Code.
- Experience with Helm, GitOps workflows, and Kubernetes automation.
- Strong scripting and automation skills using Python, Go, or Bash.
- Expertise in designing scalable CI/CD pipelines using GitHub Actions and Jenkins.
- Strong understanding of SRE principles including SLOs, SLIs, and error budgets.
- Experience with observability and monitoring tools such as Prometheus, Grafana, and Datadog.
- Hands-on exposure to alerting systems, incident management, on-call operations, and production support.
- B.Tech in Computer Science or related engineering fields.
- Candidates from strong B2B SaaS product companies with large-scale distributed systems will be preferred.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1636312