Posted on: 16/04/2026
Description :
- Strong hands-on experience with Azure Databricks and Apache Spark
- Proficient in PySpark for large-scale data processing
- Good understanding of Spark internals execution plans, partitions, shuffles
- Experience with Azure Data Lake Storage (ADLS Gen2)
- Knowledge of Azure Data Factory (ADF) for orchestration
- Experience with Delta Lake, Parquet, and other columnar formats
- Strong Python scripting skills
- Intermediate to advanced SQL (joins, subqueries, CTEs, window functions)
- Experience in performance tuning, cost optimization, and debugging in Databricks
- Exposure to CI/CD, Git, and DevOps practices is a plus
Responsibilities :
- Design, develop, and optimize data pipelines using Azure Databricks
- Implement ETL/ELT workflows using PySpark and SQL
- Work with structured and semi-structured data at scale
- Optimize Spark jobs using partitioning, caching, and efficient joins
- Implement Delta Lake best practices (ACID, schema evolution, time travel)
- Collaborate with cloud, analytics, and business teams
- Ensure data quality, reliability, and performance in production pipelines
- Monitor jobs, troubleshoot failures, and perform root cause analysis
- Follow coding standards and contribute to documentation and knowledge sharing
Soft Skills :
- Excellent communication
- Team collaboration
- Documentation and knowledge sharing
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1629088