HamburgerMenu
hirist

Job Description

About the Role :

We are looking for a Senior Data Engineer who is equally comfortable with production data pipelines and with PyTorch. You will own the data layer that our machine learning and agentic AI systems depend on ingestion, transformation, feature engineering, training data curation, and the infrastructure that gets models from a notebook into production.

This is a hands-on engineering role, not a supervisory one. You will write code every day, work directly with client engineering teams, and be accountable for the reliability and cost of what you build.

Roles and Responsibilities :

- Design, build and operate batch and streaming data pipelines that feed model training and inference at production scale.

- Build and optimise PyTorch training and inference workflows custom Datasets and DataLoaders, distributed training (DDP/FSDP), mixed precision, checkpointing and reproducibility.

- Own feature engineering and feature store design; ensure training/serving parity and prevent data leakage.

- Curate, version and validate training datasets including labelling workflows, data quality checks and drift detection.

- Deploy and serve models (TorchServe, ONNX Runtime, Triton or equivalent) with sensible latency, throughput and cost characteristics.

- Build data infrastructure for retrieval-augmented and agentic systems: embedding pipelines, vector stores, chunking strategies and evaluation datasets.

- Instrument everything pipeline observability, lineage, data contracts, model performance monitoring and alerting.

- Partner with frontend, backend and platform engineers to ship end-to-end AI features, and with client stakeholders to translate ambiguous requirements into a data design.

- Review code, raise the engineering bar, and mentor mid-level engineers on the team.

Must-Have Qualifications :

- 5+ years of professional experience in data engineering, ML engineering or a closely related discipline.

- Strong, production-grade Python. You write tested, typed, maintainable code not just scripts.

- Hands-on PyTorch experience: building and training models, writing custom data loading, debugging training runs, and taking at least one model to production.

- Deep SQL and solid data modelling fundamentals (dimensional modelling, partitioning, indexing, query optimisation).

- Production experience with a distributed processing engine Apache Spark, Ray, Dask or equivalent.

- Orchestration experience with Airflow, Dagster, Prefect or similar, including backfills, idempotency and failure recovery.

- Cloud data platform experience on AWS, GCP or Azure (S3/GCS, Glue/Dataproc, EMR, Redshift/BigQuery/Snowflake, Databricks or comparable).

- Comfortable with Docker, Git-based workflows and CI/CD; able to containerise and ship your own work.

- Clear written and spoken English you will be writing design docs and talking to enterprise clients.

Good to Have :

- Experience with LLM and agentic systems embeddings, vector databases (pgvector, Pinecone, Weaviate, Qdrant), RAG pipelines, or evaluation harnesses for LLM outputs.

- Fine-tuning or parameter-efficient tuning experience (LoRA, QLoRA) and familiarity with Hugging Face Transformers.

- MLOps tooling : MLflow, Weights & Biases, Kubeflow, SageMaker or Vertex AI.

- Kubernetes, Terraform or other infrastructure-as-code experience.

- Streaming systems: Kafka, Kinesis, Flink or Spark Structured Streaming.

- Lakehouse formats Delta Lake, Apache Iceberg or Hudi.

- dbt and modern analytics engineering practice.

- GPU performance tuning, quantisation, or inference cost optimisation.

- Open-source contributions or published technical writing.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...