Posted on: 31/08/2026
Job Description:
We're looking for a skilled Data Engineer to design, build, and maintain scalable data pipelines that power analytics, reporting, and machine learning initiatives. You'll work extensively with PySpark for large-scale data processing and the AWS ecosystem to build reliable, cloud-native data infrastructure.
Key Responsibilities:
- Design, develop, and maintain ETL/ELT pipelines using PySpark for batch and streaming data processing.
- Build and manage data lake and data warehouse solutions on AWS (S3, Redshift, Glue, EMR, Athena, Lake Formation).
- Develop and orchestrate workflows using AWS Step Functions, Apache Airflow, or AWS Glue Workflows.
- Optimize Spark jobs for performance, cost, and scalability (partitioning, caching, cluster tuning).
- Ingest data from multiple sources (APIs, databases, flat files, streaming platforms like Kafka/Kinesis).
- Implement data quality checks, validation frameworks, and monitoring/alerting for pipeline health.
- Collaborate with data analysts, data scientists, and business stakeholders to understand data requirements.
- Design and maintain data models (star/snowflake schemas) for analytics use cases.
- Write clean, well-documented, testable code following engineering best practices (CI/CD, version control).
- Ensure data security, governance, and compliance (IAM policies, encryption, access controls).
- Troubleshoot and resolve production data pipeline issues.
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1667232