HamburgerMenu
hirist

Job Description

Role Overview :

We are looking for a highly skilled Lead Data Engineer with 8 - 15 years of experience in designing, developing, and architecting scalable data engineering solutions.

The ideal candidate should have strong hands-on expertise in Advanced Python, Databricks, Unity Catalog, Delta Live Tables (DLT), dbt, Medallion Architecture, data pipelines, SQL, and cloud data platforms.

This is a hands-on technical leadership role involving solution design, architecture, development, code reviews, performance optimization, technical decision-making, and mentoring.

Tech Stack :

- Advanced Python, Databricks, Unity Catalog, Delta Live Tables (DLT), dbt, Medallion Architecture, SQL, Cloud Data Platforms.

Key Responsibilities :

Architecture & Solution Design :

- Design scalable, secure, and high-performance data engineering solutions using Databricks and modern cloud technologies.

- Translate business requirements into end-to-end data architecture and technical solutions.

- Design and implement Medallion Architecture using Bronze, Silver, and Gold layers.

- Define data ingestion, transformation, storage, governance, and consumption strategies.

- Create architecture diagrams, data flow diagrams, and technical documentation.

- Drive technical decisions related to scalability, performance, reliability, cost, and maintainability.

Advanced Python Development :

- Develop production-grade data engineering applications using Advanced Python.

- Build reusable Python frameworks, libraries, utilities, and data processing components.

- Implement complex data transformations, validations, exception handling, logging, and monitoring.

- Apply strong Python concepts including OOP, modular programming, design patterns, testing, packaging, and performance optimization.

- Develop scalable Python-based ETL/ELT pipelines and automation frameworks.

- Conduct code reviews and establish Python coding standards and best practices.

Databricks Engineering :

- Design and develop scalable data pipelines using Databricks Products, Python/PySpark, SQL, and Delta Lake.

- Work with Databricks Workflows, notebooks, compute, and job scheduling.

- Develop batch and near-real-time data processing pipelines.

- Optimize Spark jobs, SQL queries, cluster configurations, partitions, and data processing workloads.

- Implement Delta Lake capabilities including schema enforcement, schema evolution, partitioning, and optimization.

Unity Catalog :

- Implement Unity Catalog for centralized data governance and access management.

- Design catalogs, schemas, tables, views, external locations, and access controls.

- Implement role-based and fine-grained data access.

- Support data lineage, auditing, discovery, and governance requirements.

DLT / Lakeflow :

- Design and develop reliable pipelines using Delta Live Tables (DLT) and modern Databricks declarative pipeline capabilities.

- Develop batch and streaming data pipelines with appropriate data quality checks.

- Implement data quality expectations and validation rules.

- Monitor pipeline health, failures, performance, and data quality.

dbt :

- Develop dbt-based transformation frameworks integrated with Databricks.

- Create reusable models, macros, tests, sources, and documentation.

- Implement incremental models and dependency management.

- Establish data quality and testing standards using dbt.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...