Posted on: 22/09/2026
Cloud / Platform Engineer - HPC & AI Infrastructure
Location - Bangalore
About the Team :
The Software Platform team builds the cloud-native infrastructure that powers HPC and AI workloads at scale across AWS environments.
We develop internal platforms, CLI tools, Kubernetes operators, CI/CD automation, and observability systems that help engineering teams deploy, manage, and monitor complex distributed systems efficiently. You'll work across the full infrastructure lifecycle, from platform development and automation to deployment, reliability, and optimization.
What You'll Do :
- Design and develop scalable cloud infrastructure for HPC and AI workloads.
- Build and maintain cloud-native platforms, automation, and developer infrastructure.
- Develop and optimize CI/CD pipelines and automated deployment workflows.
- Work with Kubernetes, Docker, GitOps, and Infrastructure as Code (IaC).
- Build and maintain Kubernetes operators/controllers and platform automation.
- Implement infrastructure observability using OpenTelemetry, Prometheus, and Grafana.
- Design, debug, and optimize complex distributed systems.
- Collaborate with engineering teams across different functions and geographies.
What We're Looking For :
- 3+ years of hands-on experience in cloud/platform/infrastructure engineering.
- Strong programming skills in Python, Go, or Rust.
- Strong experience with Linux environments.
- Hands-on experience with AWS, Azure, or GCP.
Tech Stack :
- Kubernetes (Must)
- Ansible (Must)
- Docker / Containers
- GitOps
- Infrastructure as Code (IaC)
- CI/CD (GitHub Actions preferred)
- Prometheus, Grafana, OpenTelemetry
Ways to Stand Out :
- Experience building Kubernetes operators/controllers.
- Background in HPC, AI infrastructure, GPU computing, or distributed AI systems.
- Experience working on high-performance, large-scale cloud infrastructure.
- Strong experience with cloud-native platform engineering and automation.
Why This Role?
This is an opportunity to work beyond traditional DevOps and build the platform layer that enables high-performance AI and HPC systems to run at scale. You'll have ownership across infrastructure, automation, Kubernetes, observability, and cloud platforms while working on challenging distributed systems problems.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Cloud Computing
Job Code
1673395