Posted on: 22/04/2026
Description :
Role : GenAI Data ETL Engineer
Experience : 8 - 12 Years
Location : Gurgaon / Hyderabad
Function : Software Engineering Big Data / DWH / ETL
Role Overview :
We are seeking a GenAI Data ETL Engineer to design, build, and manage data pipelines that power LLM-based applications, copilots, and intelligent automation.
This role sits at the intersection of data engineering and generative AI, focusing on transforming complex enterprise data into high-quality inputs for retrieval-augmented generation (RAG) and advanced analytics. You will collaborate closely with GenAI engineers, platform teams, and business stakeholders to deliver scalable, secure, and future-ready data solutions.
Key Responsibilities :
1. GenAI / RAG Data Pipeline Development :
- Design and maintain ETL/ELT pipelines for structured and unstructured data sources (databases, documents, logs, APIs, SaaS tools)
- Transform data using techniques like chunking, enrichment, normalization, deduplication, and PII redaction
- Build semantic data models aligned with LLM consumption (entities, relationships, knowledge domains)
- Optimize pipelines for performance, scalability, and cost (CDC, partitioning, caching, incremental loads)
- Implement data quality checks tailored to GenAI use cases (freshness, coverage, retrieval accuracy)
2. LLM & Integration Engineering :
- Develop integrations across enterprise systems (CRM, ERP, ITSM, knowledge bases, collaboration tools)
- Enable seamless data flow into LLM orchestration frameworks (RAG pipelines, agents, workflows)
- Build logging and feedback systems for prompts, responses, and retrieval traces
- Ensure data security, governance, and compliance (access control, masking, auditability)
- Define APIs, schemas, and SLAs for reliable GenAI data consumption
3. Operations, Monitoring & Documentation :
- Implement orchestration and scheduling (Airflow, Prefect, Dagster, cloud-native tools)
- Establish observability for pipelines and retrieval systems (health, freshness, coverage)
- Troubleshoot and resolve pipeline/data issues with root-cause analysis
- Maintain documentation for data lineage, schemas, and workflows
- Collaborate with governance teams for metadata, standards, and compliance
Required Skills & Qualifications :
- Bachelors degree in Computer Science, Information Systems, or related field
- 5+ years of experience in data engineering, ETL/ELT, or data integration
- Strong SQL skills (joins, window functions, performance optimization)
- Hands-on experience with data pipeline frameworks (dbt, Airflow, Prefect, Dagster, etc.)
- Experience with cloud data platforms (Snowflake, BigQuery, Redshift, Synapse, etc.)
- Proficiency in Python (preferred) or Java/Scala for data workflows
- Experience working with APIs, JSON, CSV, and integration patterns
- Solid understanding of data modeling (relational, denormalization, CDC, event-driven ingestion)
- Familiarity with Git and standard software development practices
- Exposure to GenAI/LLM concepts through projects or production use cases
Preferred Skills :
- Experience with RAG, semantic search, and document intelligence
- Hands-on experience with vector databases (Pinecone, Weaviate, pgvector, Elasticsearch, etc.)
- Experience with orchestration tools (Airflow, Prefect, Dagster, Azure Data Factory, AWS Glue)
- Knowledge of streaming platforms (Kafka, Kinesis, Pub/Sub, EventBridge)
- Experience with data quality and observability tools (Great Expectations, Monte Carlo, Soda)
- Familiarity with cloud platforms (AWS, Azure, GCP) and their data/AI services
- Understanding of data security and compliance (encryption, access control, PII handling)
Nice to Have :
- Experience collaborating with ML/GenAI teams (feature pipelines, evaluation datasets, MLOps)
- Exposure to BI/analytics tools (Power BI, Tableau, Looker)
- Experience with data catalogs, lineage tools, or knowledge graphs
Did you find something suspicious?