HamburgerMenu
hirist

Data Engineer - Apache Spark

Lancesoft India Pvt Ltd
4 - 6 Years
Bangalore

Posted on: 19/08/2026

Job Description

Role Overview:

We are looking for an experienced Data Engineer with strong hands-on expertise in Databricks, PySpark, Python, SQL, Delta Lake, data modelling, and modern data engineering practices. The ideal candidate will be responsible for developing, maintaining, and optimizing scalable data pipelines and data platforms. The role requires hands-on experience with Databricks platform capabilities, Unity Catalog, Delta Live Tables, Auto Loader, performance optimization, monitoring, and production support.

Key Responsibilities:

Data Engineering:

- Design, develop, test, deploy, and maintain scalable and reliable data pipelines.

- Build high-performance batch and streaming data pipelines using Databricks, Apache Spark, and PySpark.

- Develop reusable and optimized data processing frameworks using Python and SQL.

- Implement data ingestion, transformation, cleansing, validation, and integration processes.

- Work with large-scale datasets and optimize data processing for performance and cost.

- Troubleshoot production issues and provide root-cause analysis and permanent fixes.

Databricks:

- Strong hands-on experience with Databricks.

- Implement and manage Databricks Unity Catalog for data governance, access control, and data discovery.

- Develop and maintain Delta Live Tables (DLT) pipelines.

- Implement incremental and event-driven ingestion using Databricks Auto Loader.

- Work extensively with Delta Lake tables and associated optimization techniques.

- Implement data quality, lineage, governance, and security practices within Databricks.

- Optimize Databricks workloads, clusters, jobs, notebooks, and Spark applications.

- Monitor Databricks workloads and troubleshoot performance or operational issues.

Data Modelling & Architecture:

- Design scalable and maintainable data models for analytical and operational use cases.

- Develop conceptual, logical, and physical data models.

- Apply appropriate dimensional modelling techniques such as Star Schema and Snowflake Schema where applicable.

- Design data architecture supporting scalability, reliability, performance, and maintainability.

- Define appropriate data storage, processing, integration, and consumption patterns.

- Ensure data solutions follow enterprise architecture and engineering standards.

Python:

- Strong proficiency in advanced Python.

- Develop production-grade Python applications and data engineering frameworks.

- Write reusable, modular, maintainable, and testable code.

- Implement exception handling, logging, configuration management, and automated testing.

- Develop utilities and frameworks to improve data engineering productivity.

PySpark / Apache Spark:

- Strong hands-on experience with advanced PySpark.

- Develop complex Spark transformations and data processing workflows.

- Optimize Spark jobs using appropriate partitioning, caching, joins, and data structures.

- Troubleshoot Spark performance issues and optimize resource utilization.

- Work with large-scale distributed data processing workloads.

SQL:

- Strong proficiency in advanced SQL.

- Develop complex queries, stored procedures, CTEs, window functions, aggregations, and analytical queries.

- Perform query optimization and troubleshoot database/data processing performance issues.

- Design efficient data extraction and transformation logic.

Logging & Monitoring:

- Implement effective logging, monitoring, alerting, and observability for data pipelines.

- Use Azure and Databricks monitoring services to monitor pipeline health, failures, performance, and resource utilization.

- Analyze application and pipeline logs to identify and resolve production issues.

- Establish appropriate operational dashboards and alerts.

Agile / Engineering Practices:

- Work closely with architects, developers, product owners, QA teams, and other stakeholders.

- Participate in Agile ceremonies including sprint planning, daily stand-ups, backlog refinement, reviews, and retrospectives.

- Follow engineering best practices for development, testing, deployment, documentation, and production support.

- Contribute to continuous improvement of data engineering processes and standards.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...