HamburgerMenu
hirist

AWS Data Lead - Python/PySpark

LA Consultancy
10 - 12 Years
Multiple Locations

Posted on: 12/08/2026

Job Description

Role Overview :

- BE, BTech, MCA, or equivalent degree in Computer Science, IT, or a related field with 10+ years of extensive hands-on experience designing and leading AWS data engineering solutions.

- Proven experience leading or owning on-premises to AWS data platform migrations.

- Advanced proficiency in Python and PySpark for data processing, automation, and orchestration.

- Strong experience with AWS Glue, Amazon S3, MWAA (Airflow), Step Functions, CloudWatch, and related services.

- Experience with Infrastructure as Code, preferably Terraform.

- Strong understanding of data engineering principles including ETL design, data modeling, and pipeline optimization.

- Demonstrated experience supporting enterprise, production-grade data platforms.

Key Responsibilities :

Technical Leadership & Platform Ownership :

- Serve as the technical lead and design authority for AWS data engineering initiatives across multiple enterprise platforms.

- Own architectural decisions related to scalability, reliability, security, and cost optimization of AWS data platforms.

- Define and enforce engineering standards, coding patterns, and operational best practices for cloud data pipelines.

AWS Data Platform Engineering :

- Lead the design, development, and support of cloud-native data pipelines using Amazon S3, AWS Glue (PySpark), MWAA (Apache Airflow), and AWS Step Functions.

- Drive on-premises to AWS data platform migrations, including reverse engineering of legacy ETL workflows and re-implementation using AWS-native services.

- Re-architect legacy Oracle Data Integrator (ODI)based ETL processes into scalable PySpark-based Glue jobs.

- Optimize Spark workloads for performance, memory usage, and cost efficiency in AWS Glue environments.

Data Lake, Iceberg & Architecture Design :

- Architect and implement enterprise AWS data lakes using Medallion architecture (Bronze, Silver, Gold).

- Design and manage Apache Iceberg tables to support incremental processing, schema evolution, and efficient data lake operations.

- Establish standardized ingestion, transformation, and consumption patterns across financial, mobility, and corporate datasets.

- Ensure data quality, reconciliation, lineage, and auditability across all layers of the data platform.

- Lead orchestration strategy using MWAA (Managed Workflows for Apache Airflow).

- Drive implementation of AWS security best practices, including IAM role design, least-privilege access, encryption using AWS KMS, and secrets management.

Preferred / Nice-to-Have :

- Experience with Apache Iceberg or similar data lake table formats.

- Exposure to analytics or BI platforms (e.g., ThoughtSpot, Tableau).

- Exposure to Oracle databases or legacy ETL tools (e.g., ODI).

- Familiarity with CI/CD practices for data engineering workloads.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...