HamburgerMenu
hirist

Lead/Principal Data Engineer - Python/Apache Spark

Geektrust
14 - 18 Years
Bangalore

Posted on: 03/09/2026

Job Description

What you need :

- 10+ years of experience in Data Engineering, with experience delivering enterprise-scale data platforms.

- Strong expertise in Databricks, Apache Spark (PySpark), Python, and SQL.

- Hands-on experience building cloud-native data platforms on AWS.

- Strong understanding of distributed data processing, Lakehouse architecture, ETL/ELT, and modern data engineering practices.

- Experience designing scalable data models and optimizing large-scale data pipelines.

- Strong software engineering fundamentals, including Git, CI/CD, testing, and code quality.

- Excellent communication skills with the ability to influence stakeholders and lead technical discussions.

What you would do :

Data Product Development :

- Own data products from design through production.

- Design, build, and maintain scalable data pipelines and data products using Databricks, Spark, Python, SQL, and AWS.

- Develop robust batch and streaming pipelines aligned with modern Lakehouse architecture principles.

- Build reusable ETL/ELT frameworks, curated datasets, and self-service data products for analytics, AI/ML, and operational reporting.

- Optimize pipelines for performance, scalability, reliability, and cost efficiency.

- Implement incremental processing, CDC, and metadata-driven engineering frameworks.

Solution Design and Architecture:

- Partner with business stakeholders to understand requirements and translate them into scalable technical solutions.

- Design end-to-end data architectures, reusable frameworks, and high-performance data models.

- Evaluate architectural trade-offs and influence technical direction across the data platform.

- Drive solution design from concept through production deployment.

Engineering Excellence :

- Build production-grade solutions with a strong focus on quality, testing, observability, security, and reliability.

- Implement automated data validation, monitoring, lineage, and operational best practices.

- Write clean, maintainable code and contribute to reusable frameworks, CI/CD, DataOps practices, and code reviews.

- Troubleshoot production issues and continuously improve platform performance and developer experience.

Business Partnership :

- Collaborate with business stakeholders, product owners, analysts, architects, and engineers to solve high-impact business problems.

- Translate business requirements into scalable data products and trusted datasets.

- Communicate technical concepts effectively to both technical and non-technical audiences.

- Take end-to-end ownership of solutions from discovery and design to production support.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...