HamburgerMenu
hirist

Senior Data Engineer - ML Platform

Supersourcing
5 - 10 Years
Bangalore

Posted on: 20/08/2026

Job Description

Position : Senior Data Engineer PyTorch / ML Data Platforms

Company : CodeWalnut

Employment Type : Full-Time with Supersourcing

Experience : 5+ Years

Location : Bangalore

Work Mode : Work From Office 5 Days a Week

Function : Data & AI Engineering

Role Overview :

We are looking for a hands-on Senior Data Engineer who is equally strong in production data engineering and PyTorch. You will own the data layer supporting machine learning and agentic AI systems, covering data ingestion, transformation, feature engineering, training data curation, and production ML infrastructure.

Key Responsibilities :

- Design, build, and operate batch and streaming data pipelines at production scale.

- Build and optimize PyTorch training and inference workflows, including custom Datasets, DataLoaders, DDP/FSDP, mixed precision, checkpointing, and reproducibility.

- Own feature engineering and feature store design, ensuring training/serving parity and preventing data leakage.

- Curate, version, and validate training datasets with data quality and drift monitoring.

- Deploy and serve ML models using TorchServe, ONNX Runtime, Triton, or equivalent platforms.

- Build data infrastructure for RAG and agentic AI systems, including embedding pipelines, vector stores, chunking, and evaluation datasets.

- Implement pipeline observability, lineage, data contracts, monitoring, and alerting.

- Collaborate with backend, frontend, platform, and client engineering teams to deliver end-to-end AI solutions.

- Review code and mentor mid-level engineers.

Mandatory Skills :

- 5+ years of Data Engineering / ML Engineering experience

- Strong Python development experience

- Hands-on PyTorch experience with production model deployment

- Strong SQL and Data Modelling fundamentals

- Experience with Apache Spark, Ray, Dask, or equivalent distributed processing technologies

- Experience with Airflow, Dagster, Prefect, or similar orchestration tools

- Experience with AWS / GCP / Azure data platforms

- Docker, Git, and CI/CD

- Strong communication and client-facing skills

Good to Have :

- LLM / Agentic AI experience

- RAG, embeddings, and Vector Databases

- Pinecone, Weaviate, Qdrant, pgvector

- LoRA / QLoRA and Hugging Face Transformers

- MLflow, Weights & Biases, Kubeflow, SageMaker, or Vertex AI

- Kubernetes / Terraform

- Kafka, Kinesis, Flink, or Spark Structured Streaming

- Delta Lake, Apache Iceberg, or Hudi

- GPU performance tuning and inference optimization

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...