Posted on: 07/09/2026
Key Responsibilities:
- Build unified AI/ML CI/CD pipelines to standardize model training, testing, and deployment across teams using tools such as Jenkins, GitHub Actions, and Kubeflow Pipelines.
- Develop reusable templates, environments, and platform components to accelerate onboarding and deployment of new ML and LLM projects.
- Automate infrastructure provisioning and management using Terraform, Helm, and Kubernetes, ensuring reliability, scalability, and consistency across environments.
- Implement robust monitoring and observability frameworks with Prometheus, Grafana, and the ELK Stack to track system health, model performance, and pipeline stability.
- Ensure security, compliance, and governance in ML and LLM workflows through best practices in secrets management, access control, and vulnerability scanning.
- Drive cost optimization and resource governance across public cloud environments (AWS, GCP, Azure) to ensure efficient use of infrastructure.
- Mentor and guide junior engineers on MLOps, LLMOps, and platform engineering best practices, fostering a culture of automation and reliability.
- Collaborate cross-functionally with ML engineers, data scientists, and DevOps teams to streamline the end-to-end model lifecycle.
- Continuously improve the AI/ML platform by identifying and implementing enhancements in automation, observability, and scalability.
Profile Qualifications:
- Total Experience: 5-8 years; relevant experience in MLOps, LLMOps, or related areas: 3+ years.
- Education in Computer Science, Machine Learning, or a related field.
- Strong problem-solving skills with excellent attention to detail.
- Proficiency in CI/CD pipeline development using tools such as Jenkins and GitHub Actions.
- Hands-on expertise in scripting and automation using Docker files, Shell, Python, and Terraform.
- Experience deploying applications on Kubernetes and managed Kubernetes services (e.g., EKS, AKS, GKE) through automated CI/CD workflows.
- Proven experience implementing monitoring and observability solutions using the ELK Stack, Prometheus, and Grafana for visualization and alerting.
- Practical experience with public cloud platforms, particularly Amazon Web Services (AWS) and Microsoft Azure.
- Experience using Kubeflow to deploy, scale, and manage ML workflows efficiently across multiple environments.
Nice-to-Have:
- Good knowledge of Generative AI concepts and Retrieval-Augmented Generation (RAG) techniques.
- Familiarity with agentic AI frameworks such as Langchain and Llama Index.
- Basic understanding of agentic AI principles and how they apply to LLM-based application development.
Did you find something suspicious?