Posted on: 17/09/2026
Senior Cloud Engineer - AI/ML
Experience : 5 - 8 Years
Location : Bangalore
Work Mode : Onsite
Joining : Immediate to 15 Days
Job Overview :
We are looking for an experienced Senior Cloud Engineer - AI/ML with strong hands-on expertise in Azure, Databricks, Kubernetes, AKS, ARO, Terraform, MLflow, CI/CD, and Python.
The ideal candidate will be responsible for supporting ML/AI development teams and building scalable, secure, and automated cloud infrastructure and deployment workflows. The role involves driving end-to-end automation for machine learning model deployment, MLOps, infrastructure provisioning, monitoring, and operations across Azure, Databricks, and Kubernetes platforms.
Key Responsibilities :
- Design, build, and maintain CI/CD and Continuous Training (CT) pipelines for ML models using Azure DevOps, GitHub Actions, Jenkins, or similar tools.
- Develop and manage deployment workflows for Databricks Jobs, MLflow models, APIs, and microservices.
- Deploy and operate ML workloads on Azure Kubernetes Service (AKS) and Azure Red Hat OpenShift (ARO).
- Automate cloud infrastructure provisioning and configuration using Terraform, Python, Bash, PowerShell, and Infrastructure as Code (IaC) practices.
- Implement GitOps-based deployment and infrastructure management wherever applicable.
- Manage and optimize Azure Databricks workspaces, clusters, compute resources, and related cloud services.
- Configure and maintain AKS/ARO clusters, networking, ingress, storage, secrets, and model-serving environments.
- Build scalable and reliable environments for ML model training, deployment, inference, and monitoring.
- Implement monitoring, logging, alerting, and observability for cloud and ML workloads.
- Troubleshoot infrastructure, deployment, Kubernetes, networking, and application-related issues.
- Collaborate closely with ML Engineers, Data Engineers, Data Scientists, DevOps Engineers, and Application Teams.
- Implement best practices for cloud security, identity and access management, governance, compliance, and cost optimization.
- Support production deployments and ensure high availability, scalability, and reliability of AI/ML workloads.
- Continuously improve automation, deployment processes, and operational efficiency.
Required Skills :
- 5 - 8 years of experience in Cloud Engineering, DevOps, MLOps, or related roles.
- Strong hands-on experience with Microsoft Azure.
- Strong experience with Azure Kubernetes Service (AKS).
- Experience with Azure Red Hat OpenShift (ARO).
- Strong knowledge of Azure Databricks and Databricks deployment/management.
- Hands-on experience with MLflow and ML model lifecycle management.
- Strong understanding of Kubernetes-based application and model deployments.
- Hands-on experience with Terraform / Infrastructure as Code (IaC).
- Strong scripting/programming skills in Python.
- Good knowledge of Bash and/or PowerShell.
- Experience building and managing CI/CD pipelines using Azure DevOps, GitHub Actions, Jenkins, or similar tools.
- Understanding of GitOps, containerization, Docker, Kubernetes, and automated deployments.
- Good understanding of cloud networking, security, IAM, secrets management, and distributed systems.
- Experience with monitoring and observability tools for cloud/Kubernetes environments.
Preferred Skills :
- Experience supporting AI/ML or MLOps platforms in enterprise environments.
- Exposure to Generative AI, LLMs, Agentic AI, or RAG pipelines.
- Experience deploying and managing AI/ML inference or model-serving workloads.
- Knowledge of Azure AI services and cloud-native AI architectures.
- Experience with GitOps tools such as Argo CD or Flux.
- Knowledge of Prometheus, Grafana, Azure Monitor, Log Analytics, or similar monitoring platforms.
- Experience with enterprise governance, security, compliance, and cost optimization.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
DevOps / Cloud
Job Code
1672059