Posted on: 20/08/2026
Position : Senior Data Engineer PyTorch / ML Data Platforms
Company : CodeWalnut
Employment Type : Full-Time with Supersourcing
Experience : 5+ Years
Location : Bangalore
Work Mode : Work From Office 5 Days a Week
Function : Data & AI Engineering
Role Overview :
We are looking for a hands-on Senior Data Engineer who is equally strong in production data engineering and PyTorch. You will own the data layer supporting machine learning and agentic AI systems, covering data ingestion, transformation, feature engineering, training data curation, and production ML infrastructure.
Key Responsibilities :
- Design, build, and operate batch and streaming data pipelines at production scale.
- Build and optimize PyTorch training and inference workflows, including custom Datasets, DataLoaders, DDP/FSDP, mixed precision, checkpointing, and reproducibility.
- Own feature engineering and feature store design, ensuring training/serving parity and preventing data leakage.
- Curate, version, and validate training datasets with data quality and drift monitoring.
- Deploy and serve ML models using TorchServe, ONNX Runtime, Triton, or equivalent platforms.
- Build data infrastructure for RAG and agentic AI systems, including embedding pipelines, vector stores, chunking, and evaluation datasets.
- Implement pipeline observability, lineage, data contracts, monitoring, and alerting.
- Collaborate with backend, frontend, platform, and client engineering teams to deliver end-to-end AI solutions.
- Review code and mentor mid-level engineers.
Mandatory Skills :
- 5+ years of Data Engineering / ML Engineering experience
- Strong Python development experience
- Hands-on PyTorch experience with production model deployment
- Strong SQL and Data Modelling fundamentals
- Experience with Apache Spark, Ray, Dask, or equivalent distributed processing technologies
- Experience with Airflow, Dagster, Prefect, or similar orchestration tools
- Experience with AWS / GCP / Azure data platforms
- Docker, Git, and CI/CD
- Strong communication and client-facing skills
Good to Have :
- LLM / Agentic AI experience
- RAG, embeddings, and Vector Databases
- Pinecone, Weaviate, Qdrant, pgvector
- LoRA / QLoRA and Hugging Face Transformers
- MLflow, Weights & Biases, Kubeflow, SageMaker, or Vertex AI
- Kubernetes / Terraform
- Kafka, Kinesis, Flink, or Spark Structured Streaming
- Delta Lake, Apache Iceberg, or Hudi
- GPU performance tuning and inference optimization
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1664667