Posted on: 04/08/2026
Job Description :
We are seeking an experienced Senior Backend & MLOps Engineer to bridge the gap between AI/ML engineering, scalable backend architecture, and production machine learning infrastructure. In this role, you will build and maintain high-performance Python microservices while architecting end-to-end ML pipelines using Kubeflow and cloud-native infrastructure. You will work closely with Data Scientists and DevOps teams to productionize, monitor, and scale AI models in production.
Key Responsibilities :
Backend Engineering (Python) :
- Design, build, and maintain high-throughput, low-latency microservices using FastAPI, Flask, or Django.
- Architect production-grade APIs for AI model inference and data processing pipelines.
- Optimize database performance (PostgreSQL, Redis, MongoDB, vector databases like Pinecone/Weaviate/Milvus).
- Implement robust asynchronous task queues using Celery, RabbitMQ, or Kafka.
MLOps & Orchestration (Kubeflow) :
- Architect, deploy, and manage production ML pipelines using Kubeflow Pipelines (KFP) and Kubeflow Notebooks.
- Implement automated CI/CD for machine learning (CT/CD) including automated retrain triggers, model evaluation, and deployment.
- Standardize model serving using frameworks such as KServe, Triton Inference Server, or BentoML.
- Manage model versioning, feature stores, and experiment tracking using tools like MLflow, Feast, or Weights & Biases.
Infrastructure & Cloud :
- Work heavily with Kubernetes (K8s), Helm, and Docker to deploy and scale AI workload clusters.
- Manage cloud-native AI infrastructure across AWS, GCP, or Azure (EKS/GKE/AKS).
- Ensure high availability, security, and cost-efficiency for GPU/CPU workloads.
- Implement robust monitoring, logging, and alerting for model drift, latency, and system health using Prometheus, Grafana, and ELK stack.
Technical Qualifications:
- Experience: 5+ years of hands-on experience in software engineering, backend development, and MLOps.
- Core Language: Advanced proficiency in Python (asyncio, memory management, multi-processing, object-oriented design).
- MLOps Core: Deep hands-on experience with Kubeflow (Pipelines, KServe, Katib) in a production environment.
- Containerization & Orchestration: Strong expertise in Docker and Kubernetes (CRDs, ingress controllers, resource limits, GPU node pools).
- ML Ecosystem: Practical knowledge of ML frameworks (PyTorch, TensorFlow, Scikit-learn) and LLM deployment frameworks (vLLM, Ollama, Hugging Face ecosystem).
- Databases & Queues: Experience with SQL/NoSQL databases, Vector DBs, and event-driven architectures (Kafka/RabbitMQ/Redis).
- CI/CD: Experience setting up GitOps pipelines (ArgoCD, GitHub Actions, GitLab CI/CD) tailored for ML workflows.
Did you find something suspicious?
Posted by
Posted in
Backend Development
Functional Area
ML / DL Engineering
Job Code
1660280