Posted on: 23/09/2026
Job Description & Summary :
- Hands-on experience in PySpark development on Azure Databricks.
- End-to-end ownership of data pipeline development, from coding and version control to build, deployment, and production validation in a modern lakehouse environment.
- PySpark and Scala codebases, build Spark JARs using Maven, and execute/validate them on Azure Databricks using Spark Submit and Databricks Workflows.
Skill :
- Apache Spark on Azure Databricks
- Azure Databricks
- Pyspark
- SQL
- Scala and Hadoop.
Responsibilities :
- Design, develop, and maintain scalable batch data pipelines using Apache Spark on Azure Databricks.
- Build Spark executable JARs using Maven and manage dependencies.
- Execute and validate Spark jobs using Spark Submit on Azure Databricks with custom runtime arguments and I/O paths.
- Implement and manage Delta Lake components (Delta Tables, Delta Live Tables).
- Work with Hive Metastore and Unity Catalog for metadata, governance, and access control.
- Making code changes in Scala/PySpark projects and taking them through to production execution.
- 6+ years of overall experience in Data Engineering, with strong focus on: Apache Spark (core APIs, transformations/actions, RDD/DataFrame/Dataset APIs).
Good To Have skill sets :
- Proficiency in Pyspark and Scala for Spark development (primary language for repositories).
- Strong SQL skills for complex queries, transformations, and performance tuning.
- Solid understanding of advanced Data Engineering concepts (e.g., data modeling in lakehouse, incremental/batch patterns, CDC, SCDs, data quality frameworks).
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1673758