Posted on: 31/08/2026
Technical Lead ML Platform & MLOps (E3):
Location: Bengaluru
Experience: 6+ years
Build the Platform Behind AI :
We are looking for a Technical Lead ML Platform & MLOps to help build the next generation of infrastructure that powers machine learning.
This is not a conventional MLOps operations role. You will work on some of the hardest engineering problems at the intersection of Machine Learning, Distributed Systems, Kubernetes, GPU infrastructure, Cloud, Developer Platforms and SRE.
The platform you build will enable Data Scientists and ML Engineers to move from an idea to production faster: Code - Features - Distributed Training - Experimentation - Model Registry - Deployment - Inference - Observability.
You will have the opportunity to design and build foundational ML platform capabilities used across multiple ML teams, influence platform architecture, solve large-scale infrastructure challenges, and establish engineering standards for how ML workloads are built and operated.
We are looking for a strong hands-on engineer who enjoys building platforms, debugging complex distributed systems, and simplifying infrastructure for hundreds of ML and engineering users.
What You Will Own :
- Architect, build and evolve scalable ML Platform and MLOps capabilities for large-scale production workloads.
- Create self-service infrastructure that allows Data Scientists and ML Engineers to train, experiment, deploy and operate models without managing underlying infrastructure.
- Design reusable platform abstractions, SDKs, APIs and tooling that dramatically improve ML developer productivity.
- Build systems for reproducibility, lineage, versioning, governance and lifecycle management across ML workflows.
- Drive architecture and technology choices for ML infrastructure.
Distributed Training & GPU Infrastructure :
- Build infrastructure for large-scale distributed training across CPU and GPU clusters.
- Design and operate multi-node and multi-GPU training environments.
- Work with distributed computing frameworks such as Ray or equivalent technologies.
- Build intelligent scheduling, resource isolation and autoscaling capabilities for ML workloads.
- Improve utilization of expensive GPU infrastructure through scheduling, workload optimization and capacity management.
- Design fault-tolerant training systems with checkpointing, retries and recovery mechanisms.
ML Workflow & Training Platform :
- Build a world-class training platform supporting the complete ML development lifecycle.
- Design scalable orchestration for training, feature engineering, validation and deployment workflows.
- Build reusable workflow components using technologies such as Airflow, Ray and Kubernetes.
- Improve scheduling, dependency management, execution isolation and reliability for thousands of ML workloads.
- Enable experimentation across different compute environments without exposing infrastructure complexity to users.
MLOps & Model Lifecycle:
- Build end-to-end ML lifecycle capabilities covering: Experimentation - Training - Validation - Model Registry - Deployment - Monitoring - Retraining.
- Build experiment tracking and model management using MLflow or equivalent technologies.
- Enable reliable model versioning, approval, rollout and rollback.
- Build automated model validation and production-readiness workflows.
- Enable reproducible ML workflows across development, staging and production environments.
Model Serving & AI Infrastructure:
- Build highly scalable infrastructure for real-time, batch and asynchronous model inference.
- Design model-serving platforms running on Kubernetes and GPU infrastructure.
- Optimize serving systems for latency, throughput, availability and cost.
- Explore and adopt technologies such as Ray Serve, NVIDIA Triton, vLLM, SGLang or equivalent platforms where appropriate.
- Enable production deployment of traditional ML, deep-learning and emerging AI/LLM workloads.
Platform Reliability & Observability:
- Treat ML infrastructure as a production-grade distributed platform.
- Define and drive SLIs, SLOs, availability and reliability standards for ML platform services.
- Build deep observability across infrastructure, pipelines, training workloads and inference systems.
- Troubleshoot challenging production issues spanning Kubernetes, GPU workloads, distributed systems, networking, storage and ML pipelines.
- Drive root-cause analysis and systematically eliminate recurring operational issues.
- Design for high availability, fault tolerance and graceful recovery.
Cloud, Kubernetes & Infrastructure:
- Design scalable compute, networking and storage infrastructure for ML workloads.
- Build and operate ML systems on Kubernetes and cloud platforms.
- Automate infrastructure using Terraform or equivalent Infrastructure-as-Code technologies.
- Build secure, isolated and reproducible runtime environments.
- Drive infrastructure efficiency through autoscaling, workload placement and cost optimization.
ML Developer Experience:
- Build developer-facing platforms, SDKs, APIs and abstractions.
- Reduce the time required to move an ML experiment into production.
- Eliminate repetitive infrastructure work for Data Scientists and ML Engineers.
- Build standardized templates and paved roads for ML development.
- Improve debugging, discoverability and observability of ML workloads.
- Enable teams to focus on models and business problems rather than infrastructure.
Technical Leadership:
- Own architecture and technical direction for major ML Platform initiatives.
- Lead complex system-design discussions and technical reviews.
- Convert ambiguous problems into scalable platform solutions.
- Drive engineering excellence across reliability, scalability, performance and maintainability.
- Mentor engineers and raise the technical bar of the team.
- Influence architecture across ML, Data, Platform, SRE and Infrastructure teams.
- Evaluate emerging technologies and make pragmatic build-vs-buy decisions.
- Take critical systems from concept through architecture, implementation and production adoption.
What We Are Looking For:
Must Have:
- 6+ years of strong hands-on software/platform engineering experience.
- Strong programming skills in Python.
- Deep hands-on experience with Kubernetes, containers and Linux.
- Experience building or operating large-scale distributed systems or platform infrastructure.
- Strong understanding of cloud infrastructure including compute, storage and networking.
- Experience building production-grade CI/CD and automation platforms.
- Experience with workflow orchestration such as Apache Airflow or equivalent systems.
- Experience with ML lifecycle tooling such as MLflow or equivalent platforms.
- Strong understanding of ML training and deployment workflows.
- Experience with Infrastructure as Code, preferably Terraform.
- Strong debugging and production troubleshooting skills.
- Experience building systems with monitoring, logging, metrics and alerting.
- Strong fundamentals in system design, reliability and distributed computing.
Strong Differentiators:
- Ray or other distributed computing frameworks.
- Distributed or multi-node ML training.
- GPU and multi-GPU infrastructure.
- Kubernetes-based ML platforms.
- ML training platforms used by multiple teams.
- Model serving and inference infrastructure.
- GPU scheduling and utilization optimization.
- Large-scale workflow orchestration.
- ML platform developer experience.
- Infrastructure cost and performance optimization.
- Feature platforms or feature stores.
Good to Have:
- Databricks, SageMaker, Vertex AI or similar ML platforms.
- Model monitoring, data drift and automated retraining.
- NVIDIA Triton, Ray Serve, vLLM or SGLang.
- LLM training/inference and LLMOps.
- Vector databases and retrieval infrastructure.
- Model governance and lineage.
- OpenTelemetry, Grafana or similar observability ecosystems.
The Kind of Engineer Who Will Thrive Here:
- You enjoy going below the abstraction.
- You don't just know how to deploy a model using a managed service you want to understand and improve what happens underneath it.
- You are equally comfortable discussing: Python, Kubernetes, Distributed Systems, GPUs, ML Training, CI/CD, MLOps, Production Reliability.
- You can take an ambiguous infrastructure problem, design the architecture, build the critical pieces, debug it under production load and create a platform that other engineers love using.
- Most importantly, you want to build foundational infrastructure that changes how ML engineering is done at scale.
Did you find something suspicious?