Posted on: 26/06/2026
Job Description :
We are seeking an experienced and highly motivated Engineering Manager AI/ML Platform to lead the design, development, and delivery of enterprise-scale AI/ML platforms and Generative AI solutions.
The ideal candidate will have a strong background in software engineering, cloud-native architectures, MLOps, and machine learning systems, coupled with proven leadership experience managing high-performing engineering teams.
This role requires driving strategic AI initiatives, including Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), AI Agents, MLOps automation, and scalable platform engineering while collaborating closely with Product, Data Science, Architecture, and Business stakeholders.
Key Responsibilities :
- Lead, mentor, and grow a team of AI/ML engineers, platform engineers, and software developers.
- Define engineering best practices, coding standards, and architectural guidelines.
- Conduct technical reviews, performance evaluations, and career development planning.
- Foster a culture of innovation, ownership, collaboration, and continuous learning.
- Drive Agile development processes and ensure timely delivery of business-critical initiatives.
- Design, develop, and maintain scalable AI/ML platforms supporting model development, training, deployment, and monitoring.
- Build reusable frameworks and platform capabilities for machine learning lifecycle management.
- Establish standardized workflows for model experimentation, validation, deployment, and governance.
- Develop enterprise-grade solutions supporting predictive analytics, machine learning, and Generative AI use cases.
- Lead implementation of GenAI applications leveraging Large Language Models (LLMs).
- Design and develop Retrieval-Augmented Generation (RAG) pipelines for enterprise knowledge systems.
- Build AI Agent architectures using frameworks such as LangChain, LangGraph, CrewAI, or similar technologies.
- Optimize prompt engineering, model orchestration, and inference performance.
- Evaluate and integrate foundation models from OpenAI, Anthropic, Meta, Google, and open-source ecosystems.
- Implement MLOps best practices for continuous training, deployment, monitoring, and retraining.
- Build automated ML pipelines using MLflow, Kubeflow, Airflow, or equivalent platforms.
- Establish model versioning, experiment tracking, feature management, and governance frameworks.
- Implement monitoring solutions for model drift, performance degradation, and operational reliability.
- Ensure reproducibility and scalability of machine learning workloads.
- Design cloud-native architectures on AWS, Azure, or Google Cloud Platform.
- Develop containerized applications using Docker and Kubernetes.
- Build scalable microservices-based systems supporting AI/ML workloads.
- Implement event-driven architectures leveraging Kafka and messaging platforms.
- Ensure high availability, fault tolerance, security, and performance optimization.
- Partner with Product Managers to define AI product roadmaps and technical strategies.
- Collaborate with Data Scientists to operationalize machine learning models.
- Work closely with Enterprise Architects to align platform capabilities with organizational objectives.
- Present technical solutions and platform strategies to senior leadership and stakeholders.
- Drive adoption of AI/ML best practices across engineering teams.
Required Skills & Technical Expertise :
1. Programming Languages :
- Strong proficiency in Python and Java.
- Experience with REST APIs and backend application development.
2. Generative AI & Machine Learning :
- Large Language Models (LLMs), Generative AI (GenAI), Retrieval-Augmented Generation (RAG), Prompt Engineering, LangChain / LangGraph, AI Agents, Fine-tuning and model optimization, and Vector databases (Pinecone, Weaviate, ChromaDB, FAISS).
3. MLOps & ML Platforms :
- MLflow, Kubeflow, Airflow, Model Monitoring, Experiment Tracking, Feature Stores, and CI/CD for ML workflows.
4. Cloud Technologies :
- AWS (SageMaker, EKS, Lambda, Bedrock), Microsoft Azure (Azure ML, AKS, OpenAI Services), and Google Cloud Platform (Vertex AI, GKE).
5. Containerization & Orchestration :
- Docker, Kubernetes, Helm, and Container Security.
6. Streaming & Messaging :
- Apache Kafka, Event-driven Architecture, and Message Queues.
7. DevOps & Automation :
- CI/CD Pipelines, GitHub Actions, Jenkins, GitLab CI/CD, and Infrastructure as Code (Terraform preferred).
8. Databases :
- PostgreSQL, MySQL, MongoDB, Redis, and Vector Databases.
Qualifications :
- Bachelor's or Master's degree in Computer Science, Engineering, Artificial Intelligence, Data Science, or related field.
- 9- 15 years of overall software engineering experience.
- 3+ years of experience leading engineering teams.
- Strong experience building enterprise AI/ML platforms and cloud-native applications.
- Proven expertise in MLOps, Generative AI, and Large Language Model-based solutions.
- Experience delivering scalable, production-grade machine learning systems.
- Strong understanding of distributed systems and microservices architecture.
- Experience with OpenAI, Anthropic, Gemini, Llama, Mistral, or similar foundation models.
- Exposure to AI governance, Responsible AI, and model compliance frameworks.
- Experience in enterprise AI transformation initiatives.
- Cloud certifications (AWS, Azure, GCP) preferred.
- Kubernetes or MLOps certifications are a plus.
Did you find something suspicious?