HamburgerMenu
hirist

Job Description

Total Experience : 7-10 years

Key Responsibilities :

- Design, build, and operate scalable cloud infrastructure and platform services for production workloads.

- Own production reliability, availability, performance, incident management, and operational excellence.

- Build and optimize CI/CD pipelines for application, infrastructure, and ML/AI deployments.

- Manage containerized workloads using Docker and Kubernetes.

- Automate infrastructure provisioning, configuration, deployment, and operational processes.

- Support microservices-based architectures and distributed systems across enterprise environments.

- Develop automation and platform tooling using Python and languages such as Node.js or Java.

- Implement MLOps workflows covering model packaging, artifact management, deployment, inference operations, versioning, and rollback.

- Operate enterprise AI/data platforms such as Databricks, MLflow, Airflow, or equivalent technologies.

- Establish robust observability, logging, alerting, tracing, and performance monitoring for distributed applications and AI workloads.

- Implement secure secrets management, access controls, auditability, and platform governance.

- Work with development, data science, ML, security, and product teams to enable reliable deployment of applications and AI solutions.

- Manage and troubleshoot complex production issues and drive root-cause analysis and preventive actions.

- Work with Apigee X or similar enterprise API gateway platforms to support secure and scalable API management.

Required Qualifications :

- 7+ years of experience in DevOps, Platform Engineering, Cloud Infrastructure, or Site Reliability Engineering.

- Strong production operations and reliability ownership.

- Hands-on experience with cloud platforms, microservices, CI/CD, containers, automation, and infrastructure architecture.

Strong expertise in :

- Docker

- Kubernetes

- Linux Administration

- Git

- System Administration

- Strong programming experience in Python and at least one additional language such as Node.js or Java.

- Hands-on experience with MLOps and ML deployment workflows.

- Experience with Databricks, MLflow, Airflow, or similar enterprise AI/data platforms.

- Strong understanding of observability, logging, alerting, monitoring, and distributed-system performance.

- Experience with secrets management, access control, security, and audit requirements.

- Working knowledge of Apigee X or equivalent API gateway technologies.

- Strong troubleshooting, problem-solving, and communication skills.

Preferred Candidate Profile :

- Experience supporting AI/ML and data-intensive production environments.

- Strong understanding of cloud-native and Kubernetes-based architectures.

- Experience building internal developer platforms or reusable infrastructure capabilities.

- Strong automation mindset with a focus on reducing manual operational effort.

- Comfortable working in high-availability, enterprise-scale production environments.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...