Posted on: 08/09/2026
Total Experience : 7-10 years
Key Responsibilities :
- Design, build, and operate scalable cloud infrastructure and platform services for production workloads.
- Own production reliability, availability, performance, incident management, and operational excellence.
- Build and optimize CI/CD pipelines for application, infrastructure, and ML/AI deployments.
- Manage containerized workloads using Docker and Kubernetes.
- Automate infrastructure provisioning, configuration, deployment, and operational processes.
- Support microservices-based architectures and distributed systems across enterprise environments.
- Develop automation and platform tooling using Python and languages such as Node.js or Java.
- Implement MLOps workflows covering model packaging, artifact management, deployment, inference operations, versioning, and rollback.
- Operate enterprise AI/data platforms such as Databricks, MLflow, Airflow, or equivalent technologies.
- Establish robust observability, logging, alerting, tracing, and performance monitoring for distributed applications and AI workloads.
- Implement secure secrets management, access controls, auditability, and platform governance.
- Work with development, data science, ML, security, and product teams to enable reliable deployment of applications and AI solutions.
- Manage and troubleshoot complex production issues and drive root-cause analysis and preventive actions.
- Work with Apigee X or similar enterprise API gateway platforms to support secure and scalable API management.
Required Qualifications :
- 7+ years of experience in DevOps, Platform Engineering, Cloud Infrastructure, or Site Reliability Engineering.
- Strong production operations and reliability ownership.
- Hands-on experience with cloud platforms, microservices, CI/CD, containers, automation, and infrastructure architecture.
Strong expertise in :
- Docker
- Kubernetes
- Linux Administration
- Git
- System Administration
- Strong programming experience in Python and at least one additional language such as Node.js or Java.
- Hands-on experience with MLOps and ML deployment workflows.
- Experience with Databricks, MLflow, Airflow, or similar enterprise AI/data platforms.
- Strong understanding of observability, logging, alerting, monitoring, and distributed-system performance.
- Experience with secrets management, access control, security, and audit requirements.
- Working knowledge of Apigee X or equivalent API gateway technologies.
- Strong troubleshooting, problem-solving, and communication skills.
Preferred Candidate Profile :
- Experience supporting AI/ML and data-intensive production environments.
- Strong understanding of cloud-native and Kubernetes-based architectures.
- Experience building internal developer platforms or reusable infrastructure capabilities.
- Strong automation mindset with a focus on reducing manual operational effort.
- Comfortable working in high-availability, enterprise-scale production environments.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
ML / DL Engineering
Job Code
1669709