Posted on: 08/09/2026
Job Overview:
We are seeking a highly motivated and strategic-minded Cloud Engineer to join our dynamic team. You will lead our reliability strategy, build self-healing architecture, eliminate operational toil through software engineering, and mentor a high-performing team of SRE and Devops engineers.
Responsibilities:
Architecture & Reliability Ownership:
- Design multi-AZ and multi-region AWS architectures for zero-downtime releases.
- Establish and enforce SLIs, SLOs, and error budgets alongside product managers.
Infrastructure as Code (IaC):
- Standardize scalable cloud infrastructure using Terraform or AWS CDK.
- Ensure state isolation, modularity, and automated deployment pipelines.
Observability & Monitoring:
- Drive signal-to-noise alerting and build observability pipelines using Prometheus/Grafana, APM, AWS CloudWatch, or OpenTelemetry.
Incident Leadership & Postmortems:
- Serve as Incident Commander for high-severity issues.
- Lead blameless post-mortems and convert system failures into concrete engineering work.
Automation & Platform Tooling:
- Treat operations as a software problem by writing custom tooling in Bash or Python to eliminate manual operational toil.
Developer Experience & CI/CD:
- Maintain robust CI/CD pipelines (GitHub Actions, Tekton ArgoCD) to empower product developers with safe self-service deployments.
FinOps & Security:
- Partner with finance and security to optimize AWS spending (Savings Plans, right-sizing) while ensuring IAM policies, KMS encryption, and VPC configurations meet compliance standards.
Mentorship & On-Call Management:
- Set sustainable on-call rotation policies to prevent burnout and mentor junior and mid-level SREs.
Qualifications:
- Experience: 6 - 10 years in DevOps, Platform Engineering, or Infrastructure.
- AWS Mastery: Deep hands-on experience across core AWS services (EKS, Lambda, RDS S3, Route 53, Transit Gateway).
- Kubernetes Expertise: Experience managing production EKS clusters, including workload scaling, ingress control, service meshes, and GitOps deployments (ArgoCD/Flux).
- Software Development: Strong coding skills in Bash or Python to build CLI tools, custom controllers, and API integrations.
- IaC Proficiency: Proven production usage of Terraform or AWS CDK.
- Observability: Track record of building metrics, logs, and trace pipelines from the ground up using modern observability stacks.
- Relevant certifications (e.g., AWS Certified DevOps Engineer, Certified Kubernetes Administrator) are a plus.
If you are passionate about AWS cloud technologies, databases, and automation, and want to be part of a dynamic team, apply now!
Experience Range : 6 - 10 years
Educational Qualifications : B. Tech/B. E, and M. Tech
Skills Required : AWS, EKS, Argocd, Flux cd, Bash, Python, Terraforms, Kubernetes
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1669689