Posted on: 07/09/2026



Core Responsibilities :
- GenAI Data Pipelines : Build ingestion pipelines for unstructured data (documents, text, media) and optimize advanced chunking and embedding workflows.
- Vector Storage : Deploy and manage vector databases (Pinecone, Qdrant, Weaviate, pgvector) for similarity search and low-latency retrieval.
- RAG & Agent Support : Engineer context-retrieval pipelines and APIs to power RAG systems, model fine-tuning datasets, and AI agent workflows.
- Data Platform & Streaming : Design batch (dbt, Airflow) and streaming (Kafka, Spark) pipelines on cloud lakehouses (Databricks, Snowflake).
- Governance & Quality : Ensure PII masking, data lineage, and security compliance before data reaches AI models.
Key Qualifications :
- Experience : 5+ years in Data Engineering, with hands-on experience supporting GenAI/LLM pipelines or NLP workloads.
- Languages : Expert in Python and SQL.
- Vector & AI Tools : Experience with vector databases and AI orchestration frameworks (LlamaIndex, LangChain).
- Data Stack : Proficiency with cloud lakehouses (Databricks/Snowflake/BigQuery) and orchestration tools (dbt, Airflow).
- Engineering : Strong knowledge of Git, Docker, and CI/CD pipelines.
Technical Stack :
- Languages : Python, SQL
- Data Platforms : Snowflake, Databricks, PostgreSQL
- Vector Databases : Pinecone, Qdrant, Weaviate, pgvector
- Orchestration & Processing : dbt, Apache Airflow, Kafka, Spark
- AI Tooling : LlamaIndex, LangChain, Hugging Face
Did you find something suspicious?