HamburgerMenu
hirist

Optum - Senior AI/ML Engineer - MLOps & LLMOps Platforms

Optum Global Solutions
9 - 14 Years
Multiple Locations

Posted on: 17/09/2026

Job Description

Senior AI/ML Engineer - MLOps & LLMOps Platforms

About the Role :

We are looking for a Senior AI/ML Engineer - MLOps & LLMOps Platforms to build and operationalize scalable AI/ML platforms and production-grade machine learning and Generative AI applications.

The role is focused on MLOps, LLMOps, model deployment, AI infrastructure, automation, monitoring, and production engineering.

You will work on building reliable platforms and workflows that enable teams to develop, deploy, monitor, and manage ML models and LLM-based applications at scale.

The ideal candidate combines strong Python and software engineering skills with hands-on experience in MLOps/LLMOps, Kubernetes, cloud platforms, CI/CD, model serving, MLflow, LLMs, RAG, and production AI systems.

Key Responsibilities :

1. MLOps & LLMOps :

- Design and implement end-to-end MLOps and LLMOps pipelines for production AI/ML workloads.

- Build automated workflows for model training, validation, deployment, serving, monitoring, and lifecycle management.

- Develop CI/CD pipelines for ML models, LLM applications, and AI services.

- Implement model and application versioning, experiment tracking, testing, release management, and rollback strategies.

- Build scalable processes for deploying and managing ML models and LLM-based applications across cloud environments.

- Automate operational workflows to improve the reliability, scalability, and efficiency of AI systems.

2. AI/ML Platform Engineering :

- Build and maintain reusable platform components for AI/ML model deployment, inference, monitoring, and governance.

- Develop APIs, services, deployment frameworks, and tooling that enable engineering teams to operationalize AI solutions.

- Design infrastructure supporting real-time and batch model inference.

- Implement scalable model serving and inference architectures.

- Integrate AI/ML platforms with enterprise applications, data platforms, and cloud infrastructure.

3. GenAI & LLM Engineering :

- Productionize LLM and Generative AI applications using models from open-source and commercial ecosystems.

- Build and deploy RAG pipelines, embedding workflows, vector search, and semantic retrieval systems.

- Develop production-ready LLM applications with appropriate evaluation, monitoring, and deployment workflows.

- Work with LLM orchestration frameworks such as LangChain, LangGraph, LlamaIndex, Semantic Kernel, or equivalent.

- Optimize AI applications for latency, scalability, reliability, quality, and cost.

4. Cloud & Infrastructure :

- Deploy AI/ML workloads using Docker, Kubernetes, and cloud-native infrastructure.

- Build infrastructure and deployment automation using modern DevOps practices.

- Work with Azure, AWS, and/or Google Cloud AI/ML services.

- Implement scalable and resilient infrastructure for model serving and AI applications.

5. Monitoring & Reliability :

- Implement monitoring and observability for models, LLM applications, inference services, and AI platforms.

- Build logging, metrics, alerting, and tracing capabilities for production AI systems.

- Monitor model/application performance, latency, availability, errors, and resource utilization.

- Troubleshoot production issues and drive improvements in reliability and operational efficiency.

Required Qualifications :

- Bachelor's degree in Computer Science, Engineering, AI, Data Science, or a related technical field.

- 8+ years of experience in AI/ML engineering, MLOps, software engineering, platform engineering, or related areas.

- Strong hands-on experience building and deploying production ML/AI systems.

- Strong experience with MLOps and/or LLMOps.

- Strong programming skills in Python.

- Experience with MLflow, Kubeflow, Azure ML, AWS SageMaker, Google Vertex AI, or equivalent platforms.

- Hands-on experience with Docker, Kubernetes, CI/CD, and cloud platforms.

- Experience with model deployment, model serving, inference pipelines, and lifecycle management.

- Strong understanding of Machine Learning and AI system development.

- Hands-on experience with LLMs, RAG, embeddings, and vector databases.

- Experience developing APIs and production-grade AI services.

- Strong understanding of monitoring, observability, and production operations.

Preferred Qualifications :

- Experience building enterprise MLOps/LLMOps platforms or shared AI infrastructure.

- Experience with LangChain, LangGraph, LlamaIndex, Semantic Kernel, or similar frameworks.

- Experience with Kafka, Redis, Elasticsearch, Databricks, or Spark.

- Experience with LLM evaluation, prompt management, model monitoring, and AI observability.

- Experience deploying open-source foundation models and commercial LLM APIs.

- Experience with distributed systems and microservices.

- Knowledge of AI governance, security, and Responsible AI practices.

- Experience working in large-scale enterprise or regulated environments.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...