Posted on: 18/07/2026
Description :
We are looking for a Senior MLOps / Machine Learning Engineer with 6+ years of experience to design, build, and scale our next-generation machine learning infrastructure. In this role, you will bridge the gap between Data Science and Core Engineering, ensuring our predictive models move from experimental notebooks to high-throughput, production-grade systems seamlessly.
This isn't a role for standing up low-traffic inference endpoints; we deal with real-world scale. You will work closely with data scientists to optimize distributed training using Ray, orchestrate complex pipelines via Airflow/Composer, and manage scalable deployments on GCP. If you love deep-diving into Python optimization, writing clean Terraform code, and architecture built for high-volume data and traffic, this role is for you.
Requirements :
- Seniority : 6+ years of experience in an MLOps, DevOps, or Data Engineering role with a heavy focus on productionizing machine learning models.
- Python Mastery : Exceptional Python programming skills with a deep understanding of asynchronous programming, performance profiling, and backend frameworks like FastAPI.
- Production ML Scale : Proven track record of deploying and monitoring ML models at scale. You know how to handle high-concurrency traffic, model drifting, and resource optimization.
- Cloud & Data Proficiency : Strong expertise in the GCP ecosystem and writing complex, optimized SQL queries for large datasets.
- Tooling Agility : Comfortable working across a diverse ecosystem of package managers and frameworks, with a keen interest in adopting high-performance tools (like uv).
- Education : Bachelors or Masters degree in Computer Science, Engineering, Mathematics, or a related technical field (or equivalent practical experience).
Job Responsibilities :
- Infrastructure & Automation : Design, provision, and maintain scalable ML infrastructure on GCP using Terraform.
- Scalable Deployment : Architect and deploy high-throughput, low-latency ML inference endpoints capable of handling heavy production traffic.
- Pipeline Orchestration : Build, monitor, and optimize robust data and ML pipelines using GCP Cloud Composer / Apache Airflow.
- Distributed Computing : Implement and scale distributed training and data processing workloads using Ray.
- MLOps Evolution : Transition our current logging and metadata tracking from native GCP tools toward advanced open-source stacks (e.g., MLflow) as our infrastructure matures.
- Collaboration & Standards : Partner with Data Science teams to standardize package management (utilizing uv, Poetry, etc.) and establish best practices for code quality, containerization (Docker), and model reproducibility.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
ML / DL Engineering
Job Code
1655434