HamburgerMenu
hirist

Lead Data & Analytics Engineer - PySpark

Java R & D
10 - 15 Years
Delhi NCR

Posted on: 29/09/2026

Job Description

Designation: Lead Data and Analytics Engineer

Position Overview:

We are seeking a hands-on Senior Data and Analytics professional who will design, build, and operate the enterprise data and analytics ecosystem covering:

- Data Engineering and Data Platforms - open-source lakehouse stack (primary), with Fabric / Snowflake / Databricks as additional platforms

- Business Intelligence (BI) and Analytics delivery

- AI-ready data preparation and modeling

- Embedded and platform-native AI capabilities within analytics tools

This role is focused on doing rather than managing - the individual will personally build pipelines, write transformation code, model data, create semantic layers, and develop BI dashboards. The platform foundation is open-source first: the candidate must be comfortable building and maintaining open-source data infrastructure (Apache Spark, Delta Lake, dbt, Airflow, Trino/Presto, Great Expectations and similar). Experience with managed platforms (Microsoft Fabric, Snowflake, Databricks) is a strong add-on but secondary to open-source depth.

Required Skills and Competencies:

Technical: Must-Have:

- Open-source data engineering stack (hands-on, production-grade):

1. Apache Spark / PySpark - personally wrote and optimised Spark jobs in production

2. Delta Lake / Iceberg / Hudi - built and operated medallion-architecture lakes

3. dbt - authored transformation models, tests and documentation

4. Apache Airflow (or Prefect / Dagster) - built and maintained DAGs in production

5. Docker - containerised data platform components; basic Kubernetes familiarity

6. CI/CD for data pipelines - GitHub Actions / GitLab CI

- Strong SQL - complex transformations, window functions, query optimisation.

- Data modelling - dimensional modelling (star/snowflake schemas), semantic layer design.

- Power BI - hands-on dashboard and semantic model development.

- End-to-end ownership - personally built and maintained pipelines and BI artefacts in production, not just designed or reviewed them.

Technical - Strong Add-on:

- Microsoft Fabric - Lakehouse, Delta tables, Direct Lake, Fabric Copilot

- Snowflake - data loading, transformation, performance tuning

- Azure Databricks - Delta Live Tables, Unity Catalog, Workflows

- Apache Kafka / Flink - streaming ingestion

- Trino / Presto - federated query

- Great Expectations / Soda - open-source data quality frameworks

- OpenMetadata / DataHub / Apache Atlas - data cataloguing and lineage

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...