Posted on: 08/07/2026
Job Description :
We are looking for a Data Engineer with strong hands-on experience in building scalable data pipelines and distributed data processing systems using PySpark and Scala. The role involves working on large datasets, developing ETL/ELT workflows, and optimizing Spark-based applications for performance and reliability.
Responsibilities :
- Develop and maintain scalable data pipelines using PySpark and Scala
- Build ETL/ELT workflows for large-scale data processing
- Work with Apache Spark for batch and streaming data processing
- Optimize Spark jobs for performance, memory usage, and cost efficiency
- Design and implement data transformations and data models
- Integrate data from multiple sources into data lakes or warehouses
- Ensure data quality, reliability, and consistency across pipelines
- Troubleshoot and resolve production issues in data processing systems
Required Skills :
- Strong experience in PySpark and Scala
- Hands-on experience with Apache Spark ecosystem
- Strong SQL skills
- Experience in building ETL/ELT pipelines
- Knowledge of distributed data processing concepts
- Experience with large-scale data handling and performance tuning
Good to Have :
- Experience with cloud platforms (AWS or Azure)
- Knowledge of data lakes and file formats like Parquet
- Experience with workflow orchestration tools (Airflow, etc.)
- Basic understanding of Kafka or streaming systems
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1652210