HamburgerMenu
hirist

Lead Data Engineer - Databricks/PySpark

TalentRabbit
7 - 12 Years
Hyderabad

Posted on: 17/07/2026

Job Description

Role & responsibilities :

- Design, develop, and maintain scalable data pipelines using Databricks, PySpark, and Delta Lake.

- Build real-time and batch data ingestion pipelines from diverse operational systems using high-performance Kafka data pipelines.

- Implement data transformations that serve digital twin platforms and operational analytics.

- 2+ years of Technical Leadership Experience.

- Integrate Kafka event streams with Databricks for real-time operational state updates.

- Implement data quality checks using Delta Live Tables expectations.

- Ensure data governance compliance through Unity Catalog (lineage, access control, metadata).

- Optimize pipeline performance, reliability, and cost efficiency.

- Write clean, well-documented, and testable code following engineering best practices.

- Collaborate with ML engineers to deliver feature-engineered datasets.

- Participate in code reviews, knowledge sharing, and continuous improvement initiatives.

- Support production data systems through monitoring, troubleshooting, and incident resolution.

- Build business data warehouse solutions using Terradata for business intelligence.

Preferred candidate profile :

Our core data platform stack includes :

1. Data Platform & Lakehouse :

- Databricks as the single point of truth for all data.

- Realtime Data Pipelines implemented using Kafka for data ingestion.

- Databricks SQL for analytical queries.

- Unity Catalog for metadata management and governance.

- Terradata for data warehouse and business intelligence.

2. Stream & Event Processing :

- Apache Kafka for real-time event ingestion.

- Structured Streaming for continuous data processing.

- Delta Live Tables for declarative, quality-enforced pipelines.

3. Data Quality :

- Delta Live Tables expectations for data validation.

- Data profiling and anomaly detection.

Preferred Qualifications :

- 7+ years of hands-on data engineering experience.

- Track record of building and maintaining production-grade data pipelines.

- Experience with Delta Live Tables for declarative pipeline development.

- Experience working in agile, cross-functional teams.

- Familiarity with time-series data patterns and operational data modelling.

Highly Desirable :

- Experience building data pipelines for digital twin or simulation platforms.

- Familiarity with operational state modeling for real-time systems.

- Exposure to physics-informed or time-series ML feature engineering.

- Experience working with distributed, multidisciplinary teams.

- Exposure to industrial domains such as Manufacturing, Logistics, or Transportation is a plus.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...