Key Responsibilities :
- Design, develop, and maintain scalable data processing pipelines using Apache Spark and PySpark.
- Develop efficient data transformation and processing solutions using Python and SQL.
- Work with large volumes of structured and unstructured data to deliver high-performance data solutions.
- Optimize Spark jobs for performance, scalability, and reliability.
- Perform data cleansing, validation, transformation, and aggregation processes.
- Develop and maintain ETL/ELT workflows to support enterprise data platforms.
- Write complex SQL queries for data extraction, analysis, and optimization.
- Implement Spark optimization techniques including partitioning, caching, and performance tuning.
- Collaborate with data architects, analysts, and business teams to deliver data-driven solutions.
- Troubleshoot data pipeline issues and ensure data quality and consistency.
- Prepare technical documentation and follow data engineering best practices.
Required Skills :
- Strong programming experience in Python for data engineering applications.
- Hands-on experience with PySpark and Apache Spark ecosystem.
- Strong SQL coding skills with experience writing complex queries and optimizing SQL performance.
- Good understanding of Spark architecture, execution model, and optimization techniques.
- Experience developing scalable batch and data processing applications.
- Strong knowledge of data structures, data processing concepts, and distributed computing.
- Experience with ETL/ELT development and data pipeline implementation.
- Understanding of data warehousing concepts and data modeling.
- Experience working with large datasets and performance optimization.
- Strong analytical, debugging, and problem-solving skills.
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Big Data / Data Warehousing / ETL
Job Code
1657600