Posted on: 09/06/2026
Job Description :
About the Role :
We are looking for a highly skilled DevOps & MLOps Engineer to design, implement, and manage scalable infrastructure and machine learning operations pipelines.
The ideal candidate will have hands-on experience in cloud platforms, CI/CD automation, containerization, infrastructure as code, and ML model deployment and monitoring.
You will work closely with Data Scientists, ML Engineers, and Software Development teams to streamline the end-to-end machine learning lifecycle.
Key Responsibilities :
DevOps Responsibilities :
- Design, implement, and maintain cloud-native infrastructure on AWS, Azure, or GCP.
- Build and manage CI/CD pipelines for application and infrastructure deployments.
- Automate infrastructure provisioning using Infrastructure as Code (Terraform, CloudFormation, or similar tools).
- Manage containerized environments using Docker and Kubernetes.
- Monitor system performance, reliability, security, and cost optimization.
- Implement logging, monitoring, and alerting solutions using Prometheus, Grafana, ELK Stack, CloudWatch, or equivalent tools.
- Ensure high availability, scalability, and disaster recovery capabilities.
MLOps Responsibilities :
- Build and maintain ML pipelines for model training, validation, deployment, and monitoring.
- Automate model deployment workflows and manage model versioning.
- Implement ML lifecycle management using tools such as MLflow, Kubeflow, SageMaker, Vertex AI, or Azure ML.
- Collaborate with Data Science teams to operationalize machine learning models.
- Establish model monitoring, drift detection, and retraining mechanisms.
- Optimize model serving infrastructure for performance and scalability.
- Ensure governance, reproducibility, and compliance across ML workflows.
Required Skills & Qualifications :
- 3.55 years of experience in DevOps, Cloud Engineering, SRE, or MLOps roles.
- Strong experience with cloud platforms : AWS, Azure, or Google Cloud Platform.
- Hands-on experience with Docker and Kubernetes.
- Proficiency in CI/CD tools such as Jenkins, GitHub Actions, GitLab CI/CD, or Azure DevOps.
- Experience with Infrastructure as Code tools such as Terraform.
- Strong scripting skills in Python, Shell, or Bash.
- Experience with Linux system administration.
- Knowledge of monitoring and observability tools (Prometheus, Grafana, ELK, Datadog, Splunk, etc.).
- Experience with ML deployment frameworks and MLOps platforms.
- Understanding of model versioning, feature stores, experiment tracking, and model monitoring.
- Familiarity with Git and software development best practices.
Preferred Qualifications :
- Experience with Kubeflow, MLflow, Airflow, Argo Workflows, or SageMaker.
- Knowledge of distributed computing frameworks such as Spark.
- Exposure to Generative AI, LLM deployment, and vector databases.
- Experience with security best practices, DevSecOps, and cloud governance.
- Relevant cloud certifications (AWS, Azure, GCP, Kubernetes) are a plus.
Key Competencies :
- Strong problem-solving and troubleshooting skills.
- Excellent communication and stakeholder management abilities.
- Ability to work in cross-functional Agile teams.
- Ownership mindset with a focus on automation and operational excellence.
Nice to Have :
- Experience with LLMOps, RAG pipelines, and AI model deployment.
- Knowledge of GPU infrastructure management and optimization.
- Exposure to AI/ML observability tools and model governance frameworks.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
DevOps / Cloud
Job Code
1642890