HamburgerMenu
hirist

Capgemini - Data Engineer - Generative AI/LLM Pipelines

Capgemini Technology Services
5 - 10 Years
Multiple Locations

Posted on: 07/09/2026

showcase-imageshowcase-imageshowcase-image

Job Description

Core Responsibilities :

- GenAI Data Pipelines : Build ingestion pipelines for unstructured data (documents, text, media) and optimize advanced chunking and embedding workflows.

- Vector Storage : Deploy and manage vector databases (Pinecone, Qdrant, Weaviate, pgvector) for similarity search and low-latency retrieval.

- RAG & Agent Support : Engineer context-retrieval pipelines and APIs to power RAG systems, model fine-tuning datasets, and AI agent workflows.

- Data Platform & Streaming : Design batch (dbt, Airflow) and streaming (Kafka, Spark) pipelines on cloud lakehouses (Databricks, Snowflake).

- Governance & Quality : Ensure PII masking, data lineage, and security compliance before data reaches AI models.

Key Qualifications :

- Experience : 5+ years in Data Engineering, with hands-on experience supporting GenAI/LLM pipelines or NLP workloads.

- Languages : Expert in Python and SQL.

- Vector & AI Tools : Experience with vector databases and AI orchestration frameworks (LlamaIndex, LangChain).

- Data Stack : Proficiency with cloud lakehouses (Databricks/Snowflake/BigQuery) and orchestration tools (dbt, Airflow).

- Engineering : Strong knowledge of Git, Docker, and CI/CD pipelines.

Technical Stack :

- Languages : Python, SQL

- Data Platforms : Snowflake, Databricks, PostgreSQL

- Vector Databases : Pinecone, Qdrant, Weaviate, pgvector

- Orchestration & Processing : dbt, Apache Airflow, Kafka, Spark

- AI Tooling : LlamaIndex, LangChain, Hugging Face

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...