Posted on: 07/09/2026



Key Responsibilities :
- Multi-Cloud Architecture: Design and secure scalable ML infrastructure on AWS, Azure, and GCP using Terraform / IaC.
- Kubernetes Orchestration: Deploy and manage EKS/AKS/GKE clusters optimized for compute-heavy ML/LLM workloads, GPUs, and distributed training (Ray, Kubeflow).
- LLMOps & GenAI: Operationalize LLM pipelines including fine-tuning, RAG, vector databases, and high-performance inference (vLLM, Triton Server).
- CI/CD/CT Automation: Build automated pipelines for model training, packaging, testing, and deployment (MLflow, GitHub Actions, DVC).
- Observability & Drift Monitoring: Implement monitoring for model latency, resource usage, and data/concept drift (Prometheus, Grafana, Evidently AI).
- Cross-Functional Enablement: Work with Data Science and Platform teams to build self-service AI tools and enforce platform governance.
Preferred Candidate Profile :
- Experience : 5+ years in Platform/DevOps Engineering, with 3+ years dedicated to MLOps / AI Platform Engineering.
- Kubernetes : Deep hands-on experience with K8s, Helm, containerization (Docker), and orchestration engines.
- Python Development : Strong software development skills in Python (clean architecture, APIs via FastAPI/Flask).
- CI/CD & IaC : Production experience with Terraform and CI/CD tools (GitHub Actions, Azure DevOps, or GitLab CI).
- LLMOps Expertise : Direct experience managing LLM pipelines, vector databases, and inference optimization.
- Monitoring : Hands-on setup of operational and model-drift alerting systems.
- Multi-Cloud : Proficient in at least two major cloud platforms (AWS, Azure, GCP).
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
ML / DL Engineering
Job Code
1669054