Posted on: 29/08/2026
Relevant Experience : 4 to 6 years
Job Description :
Infrastructure & Pipeline Development :
- Design and maintain Argo Workflows / Argo Eventsbased ML and data processing pipelines across environments (dev / QA / staging / prod).
- Design and maintain WorkflowTemplates, DAGs, fan-out patterns, scheduled workflows, and event-driven triggers following GitOps best practices.
- Design and develop Python-based workflows and components for ML data processing and model inference.
- Work with Databricks and Spark-based workloads for data processing, transformation, and pipeline execution.
- Build and maintain Docker images for ML and GPU-accelerated workloads, including multi-stage builds and container optimization.
- Develop scalable inference workflows using Ray where required.
- Integrate with Azure ML and Azure cloud services for model versioning, model artifacts, training workflows, and deployment.
- Maintain CI/CD pipelines with automated testing, quality gates, and environment-specific deployment processes.
Model Deployment & Operations :
- Automate deployment, monitoring, and rollback of machine learning and deep learning models on production Kubernetes environments.
- Manage model versioning and artifact movement from cloud ML workspaces and storage systems to production workloads.
- Maintain configuration systems and reusable workflow components to ensure reproducible pipeline execution across environments.
- Monitor model and pipeline performance, data processing health, resource utilization, and infrastructure reliability.
- Troubleshoot production issues involving failed workflows, data bottlenecks, resource contention, GPU utilization, and latency regressions.
- Identify reliability and performance issues proactively and implement improvements to ML and data processing workflows.
- Collaborate with data scientists and engineering teams to improve model deployment, inference, and pipeline reliability.
Expertise and Qualifications :
Must-Have Technical Skills :
- Python : Strong programming and problem-solving skills; experience with application development, scripting, and Python packaging.
- Workflow Design : Ability to design reliable workflows with dependencies, parallel execution, retries, failure handling, scheduling, and event-driven execution.
- Argo Workflows / Argo Events : Understanding of WorkflowTemplates, DAGs, event-driven workflows, and workflow execution.
- Databricks / Data Processing : Hands-on experience with Databricks, Spark/Spark SQL, or similar large-scale data processing platforms.
- Kubernetes : Understanding of workload orchestration, resource management, namespaces, and production troubleshooting.
- Git / Version Control : Strong understanding of Git or equivalent version control systems, including branching, merging, pull requests, and maintaining code/configuration changes.
- Docker : Containerization, multi-stage builds, and optimization of ML workloads.
- CI/CD : Experience designing pipelines with automated testing, quality gates, artifact management, and deployment processes.
- Azure Cloud : Working knowledge of Azure ML, Azure Blob Storage, AKS, and Azure-based ML workflows.
- Problem Solving : Ability to troubleshoot unfamiliar technical problems and identify practical solutions.
Good-to-Have Technical Skills :
- Ray : distributed data processing and inference at scale.
- Deep Learning frameworks PyTorch, TensorFlow; model formats such as ONNX, SavedModel, and TorchScript.
- Helm / ArgoCD and GitOps-based deployment practices.
- Prometheus / Grafana and production observability.
- Experience with Azure Pipelines or similar CI/CD platforms.
- Experience with ML model monitoring, data drift, and model performance monitoring.
- Large-scale data processing experience working with large volumes of structured, unstructured, image, or video data, particularly for ML/AI workloads.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
ML / DL Engineering
Job Code
1666932