HamburgerMenu
hirist

Data Engineer - PySpark/ETL

Tranzeal
7 - 12 Years
Bangalore

Posted on: 31/08/2026

Job Description

Job Description:

We're looking for a skilled Data Engineer to design, build, and maintain scalable data pipelines that power analytics, reporting, and machine learning initiatives. You'll work extensively with PySpark for large-scale data processing and the AWS ecosystem to build reliable, cloud-native data infrastructure.

Key Responsibilities:

- Design, develop, and maintain ETL/ELT pipelines using PySpark for batch and streaming data processing.

- Build and manage data lake and data warehouse solutions on AWS (S3, Redshift, Glue, EMR, Athena, Lake Formation).

- Develop and orchestrate workflows using AWS Step Functions, Apache Airflow, or AWS Glue Workflows.

- Optimize Spark jobs for performance, cost, and scalability (partitioning, caching, cluster tuning).

- Ingest data from multiple sources (APIs, databases, flat files, streaming platforms like Kafka/Kinesis).

- Implement data quality checks, validation frameworks, and monitoring/alerting for pipeline health.

- Collaborate with data analysts, data scientists, and business stakeholders to understand data requirements.

- Design and maintain data models (star/snowflake schemas) for analytics use cases.

- Write clean, well-documented, testable code following engineering best practices (CI/CD, version control).

- Ensure data security, governance, and compliance (IAM policies, encryption, access controls).

- Troubleshoot and resolve production data pipeline issues.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...