HamburgerMenu
hirist

Job Description

Job Description :

About the Role :

We are looking for a highly skilled DevOps & MLOps Engineer to design, implement, and manage scalable infrastructure and machine learning operations pipelines.


The ideal candidate will have hands-on experience in cloud platforms, CI/CD automation, containerization, infrastructure as code, and ML model deployment and monitoring.


You will work closely with Data Scientists, ML Engineers, and Software Development teams to streamline the end-to-end machine learning lifecycle.

Key Responsibilities :

DevOps Responsibilities :

- Design, implement, and maintain cloud-native infrastructure on AWS, Azure, or GCP.

- Build and manage CI/CD pipelines for application and infrastructure deployments.

- Automate infrastructure provisioning using Infrastructure as Code (Terraform, CloudFormation, or similar tools).

- Manage containerized environments using Docker and Kubernetes.

- Monitor system performance, reliability, security, and cost optimization.

- Implement logging, monitoring, and alerting solutions using Prometheus, Grafana, ELK Stack, CloudWatch, or equivalent tools.

- Ensure high availability, scalability, and disaster recovery capabilities.

MLOps Responsibilities :

- Build and maintain ML pipelines for model training, validation, deployment, and monitoring.

- Automate model deployment workflows and manage model versioning.

- Implement ML lifecycle management using tools such as MLflow, Kubeflow, SageMaker, Vertex AI, or Azure ML.

- Collaborate with Data Science teams to operationalize machine learning models.

- Establish model monitoring, drift detection, and retraining mechanisms.

- Optimize model serving infrastructure for performance and scalability.

- Ensure governance, reproducibility, and compliance across ML workflows.

Required Skills & Qualifications :

- 3.55 years of experience in DevOps, Cloud Engineering, SRE, or MLOps roles.

- Strong experience with cloud platforms : AWS, Azure, or Google Cloud Platform.

- Hands-on experience with Docker and Kubernetes.

- Proficiency in CI/CD tools such as Jenkins, GitHub Actions, GitLab CI/CD, or Azure DevOps.

- Experience with Infrastructure as Code tools such as Terraform.

- Strong scripting skills in Python, Shell, or Bash.

- Experience with Linux system administration.

- Knowledge of monitoring and observability tools (Prometheus, Grafana, ELK, Datadog, Splunk, etc.).

- Experience with ML deployment frameworks and MLOps platforms.

- Understanding of model versioning, feature stores, experiment tracking, and model monitoring.

- Familiarity with Git and software development best practices.

Preferred Qualifications :

- Experience with Kubeflow, MLflow, Airflow, Argo Workflows, or SageMaker.

- Knowledge of distributed computing frameworks such as Spark.

- Exposure to Generative AI, LLM deployment, and vector databases.

- Experience with security best practices, DevSecOps, and cloud governance.

- Relevant cloud certifications (AWS, Azure, GCP, Kubernetes) are a plus.

Key Competencies :

- Strong problem-solving and troubleshooting skills.

- Excellent communication and stakeholder management abilities.

- Ability to work in cross-functional Agile teams.

- Ownership mindset with a focus on automation and operational excellence.

Nice to Have :

- Experience with LLMOps, RAG pipelines, and AI model deployment.

- Knowledge of GPU infrastructure management and optimization.

- Exposure to AI/ML observability tools and model governance frameworks.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...