HamburgerMenu
hirist

Platform Engineer - Cloud Infrastructure

Consulting Pandits
6 - 10 Years
Pune

Posted on: 18/09/2026

Job Description

Key Responsibilities :

- Build and operate Kubernetes clusters, with cloud-hosted control planes and AI accelerator nodes joined as workers over site-to-site connectivity.

- Register, label and taint accelerator worker nodes so that inference workloads schedule onto the correct hardware class and manage device scheduling and topology constraints.

- Plan and execute cluster and operating system upgrades: RKE2 version upgrades, RHEL patching and major-version migration, etcd backup and restore, and control-plane node replacement.

- Own cluster networking and storage end to end: CNI, ingress, DNS, load balancing, CSI drivers, persistent volume lifecycle, backup and tested disaster recovery.

- Deploy, configure and upgrade the vendor AI platform stack, which is delivered as Helm charts from an OCI registry and must be installed in a defined dependency order.

- Manage platform configuration as code: Helm values files, chart versions, namespace layout, registry pull secrets, artifact credentials and service-account key rotation.

- Manage TLS certificates and DNS for the inference API and console endpoints, including CA-issued and wildcard certificates and automated renewal.

- Operate the supporting data services the stack depends on, including operator-managed PostgreSQL, Redis queues and the bundled identity provider.

- Design and operate cloud network infrastructure: virtual networks, subnets, routing, security groups, NAT and controlled egress, with ongoing cost analysis and right-sizing.

- Own our side of IPSec connectivity into the accelerator racks, including tunnel endpoints, client-side routing and failover, and keep hybrid path latency inside inference latency budgets.

- Build and maintain Terraform modules and Ansible automation, and reconcile cluster and platform state from version control through a GitOps workflow.

- Implement cloud IAM, Kubernetes RBAC, namespace isolation, pod security standards, secrets rotation and hardening baselines, and produce evidence for security reviews.

- Deploy and operate the monitoring and logging stack, define service-level objectives and alerts tied to inference availability and latency, and track cluster and accelerator capacity.

- Support model bundle and deployment configuration changes through the platform's Kubernetes custom resources, in coordination with ML systems engineers.

- Lead incident response for cluster and platform faults, write root-cause analyses that result in a tracked change, and maintain runbooks as a deliverable of each change.

Minimum Requirements :

- Strong Linux administration on enterprise distributions, at the level of diagnosing service, storage, network, and kernel problems without escalation.

- Production Kubernetes lifecycle experience: building clusters, upgrading them and recovering them when they break. RKE2, K3s or another CNCF-certified distribution is preferred over managed-only experience.

- Helm proficiency beyond installing public charts: values management, chart versioning, multi-chart upgrade and rollback, and debugging failed releases.

- Deep hands-on experience with at least one major public cloud and working knowledge of a second, covering networking, identity and cost management.

- Terraform and Ansible at production scale, as reusable and reviewed code rather than one-off scripts.

- Networking fundamentals: routing, NAT, firewalling, DNS and TLS termination, plus the ability to debug a hybrid connectivity problem end to end.

- Working knowledge of OIDC authentication and how identity providers integrate with Kubernetes and platform applications.

- Practical experience running a Prometheus and Grafana monitoring stack and a centralised log pipeline.

- Scripting in Python and Bash, and comfort with YAML-heavy configuration.

- Strong ownership and automation instinct, clear written communication for runbooks and incident reports, and availability for a shared on-call rotation.

Preferred Requirements :

- Experience operating AI or HPC clusters, including accelerator-aware scheduling and node health management.

- Exposure to non-GPU AI accelerators and their distinct driver, runtime and scheduling models.

- Experience deploying a vendor-supplied platform product into a customer or partner environment, including handover and upgrade cycles.

- Policy-as-code tooling such as OPA, Kyverno or Sentinel, and experience with air-gapped or restricted-egress deployments.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...