HamburgerMenu
hirist

Platform Engineer - HPC & AI Infrastructure

Square Root Consulting
3 - 10 Years
rupee50-90 LPA
Bangalore

Posted on: 22/09/2026

Job Description

Cloud / Platform Engineer - HPC & AI Infrastructure

Location - Bangalore

About the Team :

The Software Platform team builds the cloud-native infrastructure that powers HPC and AI workloads at scale across AWS environments.

We develop internal platforms, CLI tools, Kubernetes operators, CI/CD automation, and observability systems that help engineering teams deploy, manage, and monitor complex distributed systems efficiently. You'll work across the full infrastructure lifecycle, from platform development and automation to deployment, reliability, and optimization.

What You'll Do :

- Design and develop scalable cloud infrastructure for HPC and AI workloads.

- Build and maintain cloud-native platforms, automation, and developer infrastructure.

- Develop and optimize CI/CD pipelines and automated deployment workflows.

- Work with Kubernetes, Docker, GitOps, and Infrastructure as Code (IaC).

- Build and maintain Kubernetes operators/controllers and platform automation.

- Implement infrastructure observability using OpenTelemetry, Prometheus, and Grafana.

- Design, debug, and optimize complex distributed systems.

- Collaborate with engineering teams across different functions and geographies.

What We're Looking For :

- 3+ years of hands-on experience in cloud/platform/infrastructure engineering.

- Strong programming skills in Python, Go, or Rust.

- Strong experience with Linux environments.

- Hands-on experience with AWS, Azure, or GCP.

Tech Stack :

- Kubernetes (Must)

- Ansible (Must)

- Docker / Containers

- GitOps

- Infrastructure as Code (IaC)

- CI/CD (GitHub Actions preferred)

- Prometheus, Grafana, OpenTelemetry

Ways to Stand Out :

- Experience building Kubernetes operators/controllers.

- Background in HPC, AI infrastructure, GPU computing, or distributed AI systems.

- Experience working on high-performance, large-scale cloud infrastructure.

- Strong experience with cloud-native platform engineering and automation.

Why This Role?

This is an opportunity to work beyond traditional DevOps and build the platform layer that enables high-performance AI and HPC systems to run at scale. You'll have ownership across infrastructure, automation, Kubernetes, observability, and cloud platforms while working on challenging distributed systems problems.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...