Posted on: 18/05/2026
Role Overview :
As a Data Engineer, you will be responsible for designing, building, and optimizing robust data platforms and pipelines across cloud and on-premises environments. The role focuses on scalable data ingestion, transformation, storage, processing, and integration to support analytics, reporting, machine learning, and AI/GenAI use cases. You will work closely with Data Scientists, ML Engineers, and business teams to ensure high-quality, governed, and reliable data is available across the enterprise.
Key Responsibilities :
- Design, build, and maintain scalable ETL/ELT pipelines for structured, semi-structured, and unstructured data
- Develop and manage batch, near real-time, and streaming data pipelines using Azure Data Factory, Spark, Kafka, Event Hubs, and Databricks
- Build and maintain data lake, lakehouse, and warehouse architectures across cloud and on-premises platforms
- Create and maintain nodes, edges, labels, properties, relationship types, and graph schemas based on business and engineering requirements.
- Support Knowledge Graph creation for engineering knowledge discovery, dependency mapping, traceability, impact analysis, and AI-powered search.
- Enable Graph RAG / Knowledge Graph + RAG patterns to improve contextual retrieval and explainable AI responses.
- Implement data ingestion frameworks for enterprise applications, APIs, files, databases, and third-party systems
- Design and optimize data models for analytics, reporting, operational, and AI-driven workloads
- Work extensively with PostgreSQL and other relational databases for data storage, transformation, performance tuning, and integration
- Build reusable data transformation frameworks using Python, SQL, and PySpark
- Ensure data quality through validation, reconciliation, exception handling, and monitoring mechanisms
- Implement metadata management, lineage, cataloging, and governance practices using Purview and enterprise data controls
- Optimize data pipelines and storage layers for performance, scalability, reliability, and cost efficiency
- Develop data orchestration workflows using Azure Data Factory, Databricks Workflows, and Airflow
- Support hybrid data integration between cloud and on-premises systems
- Implement secure data access using Azure AD, RBAC, Key Vault, encryption, and enterprise security controls
- Enable curated and governed datasets for BI, analytics, machine learning, and GenAI applications
- Support vector and embedding pipelines for advanced search and RAG use cases where required
- Collaborate with cross-functional teams to understand source systems, data requirements, and downstream consumption needs
- Maintain technical documentation for pipelines, schemas, models, workflows, and integration patterns
- Contribute to CI/CD automation, release management, and deployment of data engineering solutions using Azure DevOps or GitHub Actions
Required Skills :
- Proficiency in Python, SQL, PySpark.
- Strong working knowledge of PostgreSQL, including query tuning, indexing, partitioning, and database design
- Understanding of data security, governance, lineage, and compliance practices
- Hands-on experience with ETL pipelines and data modeling.
- Experience in building Knowledge Graphs for enterprise search, document intelligence, engineering traceability, or AI/GenAI use cases.
- Knowledge of cloud data platforms (AWS, Azure, or GCP).
- Experience with Databricks, Azure Data Factory, Synapse, Data Lake, and distributed data processing frameworks
- Experience integrating data across cloud and on-premises environments
- Familiarity with Docker, Kubernetes, CI/CD pipelines.
Preferred Skills :
- Exposure to real-time and event-driven data architecture
- Experience with Kafka, Event Hubs, or streaming data platforms
- Exposure to lakehouse architecture and modern data platform design
- Experience supporting ML, AI, or GenAI data workloads
- Familiarity with vector databases, embeddings, semantic search, or RAG-based retrieval pipelines
- Exposure to MLOps / LLMOps practices
Did you find something suspicious?
Posted by
Abhiraj B
Last Active: NA as recruiter has posted this job through third party tool.
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1636811