HamburgerMenu
hirist

Engineering Manager - MLOps

HexaCorp
8 - 10 Years
Anywhere in India/Multiple Locations

Posted on: 25/09/2026

Job Description

Purpose & Scope :

The ML Ops Engineering Manager is responsible for leading the delivery and operational excellence of ML Ops capability - the infrastructure, pipelines, and practices that take machine learning and AI models from development into reliable, governed production use.

This role manages a team of ML Ops engineers and partners closely with Data Science, Data Engineering, Platform, and Governance teams to deliver scalable, secure, and well-monitored model deployment and operations across growing AI portfolio.

The ML Ops Engineering Manager focuses on execution, engineering rigor, team leadership, and cross-functional coordination, while model strategy and prioritization remain with Data Science and AI leadership.

What you will be doing (responsibilities) :

ML Ops Delivery Leadership :

- Lead end-to-end delivery of CI/CD pipelines, model registry practices, and deployment infrastructure for ML and AI use cases.

- Drive predictable execution of the ML Ops roadmap in partnership with Data Science and Platform leadership.

- Establish and enforce engineering standards for model packaging, testing, deployment, and rollback.

- Proactively manage delivery risks, technical dependencies, and production incidents across deployed models.

Platform & Architecture Alignment :

- Ensure ML Ops solutions align with enterprise data and AI platform standards and architecture patterns.

- Partner with platform and architecture teams to design scalable, cost-effective serving and training infrastructure.

- Guide teams on appropriate use of shared compute, environments, and model infrastructure.

Monitoring, Governance & Trust :

- Embed model monitoring, drift detection, and performance alerting into pipelines as standard practice.

- Ensure model versioning, lineage, and documentation requirements are met to support auditability.

- Partner with data governance, security, and compliance teams to ensure responsible and compliant AI deployment.

Engineering Excellence & Operational Readiness :

- Drive CI/CD maturity for ML pipelines, including automated testing, staged rollouts, and controlled promotions.

- Ensure deployed models are operationally ready with monitoring, alerting, and clear incident ownership.

- Continuously improve reliability, latency, and cost-efficiency of training and inference workloads.

Stakeholder & Cross-Functional Collaboration :

- Partner with Data Science and AI leadership to translate model roadmaps into executable engineering deliverables.

- Collaborate with Data Engineering to ensure consistent, high-quality data feeds into ML pipelines.

- Communicate delivery status, risks, and trade-offs clearly to stakeholders and leadership.

People Leadership & Team Development :

- Manage, mentor, and develop a team of ML Ops engineers across experience levels.

- Set clear expectations around quality, delivery discipline, and operational ownership.

- Foster a culture of automation, documentation, and continuous improvement.

What you bring (Qualifications) :

Required :

- 8 - 10 years of experience in MLOps, ML engineering, DevOps, or platform engineering, including team or delivery leadership.

- Strong hands-on background in CI/CD, containerization, and orchestration for ML workloads.

- Experience operating model registries, monitoring tooling, and ML pipelines in enterprise environments.

- Working knowledge of cloud platforms (Azure preferred) and Databricks-based ecosystems.

- Strong stakeholder management skills across data science, engineering, platform, and governance teams.

Preferred :

- Experience supporting AI/ML programs in retail, consumer goods, or other data-intensive industries.

- Familiarity with LLM/GenAI deployment patterns and evaluation practices.

- Exposure to enterprise data governance and AI risk/compliance frameworks.

Success Measures :

- Predictable, governed delivery of ML Ops capabilities with reduced rework.

- Improved model deployment reliability, observability, and incident response.

- Increased reuse of standardized deployment patterns across model teams.

- High stakeholder confidence in the reliability and execution of the ML Ops function.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...