Posted on: 25/03/2026
Role Overview :
As a PySpark Developer, you will be instrumental in designing, developing, and maintaining robust data pipelines and solutions using PySpark within an Azure environment. Your daily activities will involve transforming raw data into actionable insights, collaborating closely with data scientists, data engineers, and business stakeholders to understand data requirements and deliver high-quality data products. Your work will directly impact the business by enabling data-driven decision-making, improving operational efficiency, and enhancing customer experiences through advanced analytics and reporting.
Key Responsibilities :
- Develop and maintain scalable data pipelines using PySpark to ingest, process, and transform large datasets from various sources for data warehousing and analytics purposes.
- Design and implement efficient data models and schemas in Azure data storage solutions (e.g., Azure Data Lake Storage, Azure Synapse Analytics) to optimize data retrieval and analysis for business users.
- Collaborate with data scientists and business analysts to understand data requirements and translate them into technical specifications and data solutions that meet their analytical needs.
- Optimize PySpark code for performance and scalability, ensuring efficient data processing and minimizing resource consumption to support growing data volumes and user demands.
- Implement data quality checks and validation processes within data pipelines to ensure data accuracy and reliability for downstream analytics and reporting.
- Troubleshoot and resolve data pipeline issues, working closely with infrastructure and operations teams to maintain data availability and system stability for critical business processes.
- Contribute to the development of data engineering best practices and standards, promoting code reusability, maintainability, and scalability across the organization to improve overall data management capabilities.
Required Skillset :
- Demonstrated proficiency in developing data pipelines and data transformation processes using PySpark, showcasing the ability to handle large datasets and complex data structures.
- Strong understanding of data warehousing concepts, data modeling techniques, and database systems, enabling the design of efficient and scalable data solutions.
- Experience working with Azure cloud services, including Azure Data Lake Storage, Azure Synapse Analytics, and Azure Data Factory, highlighting the ability to leverage cloud-based data processing and storage solutions.
- Excellent problem-solving and analytical skills, with the ability to identify and resolve data-related issues and optimize data processing performance.
- Effective communication and collaboration skills, demonstrating the ability to work with cross-functional teams and communicate technical concepts to both technical and non-technical audiences.
- Bachelor's or Master's degree in Computer Science, Data Engineering, or a related field.
- Adaptable to a remote or hybrid work environment, demonstrating the ability to collaborate effectively with distributed teams and manage time efficiently.
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1623584