HamburgerMenu
hirist

Job Description

Role Summary :

We are looking for an experienced Sr. AWS Data Engineer with 8+ years of expertise in designing, building, and optimizing large-scale cloud-based data platforms. The ideal candidate should have strong hands-on experience with AWS data services, distributed computing frameworks, and real-time streaming architectures. You will be responsible for developing scalable data pipelines, implementing modern data lake solutions, and enabling high-performance analytics using AWS and Apache Spark technologies.

Key Technical Skills :

- AWS Glue

- AWS EMR (Elastic MapReduce)

- AWS Glue ETL

- AWS Glue Data Catalog

- PySpark

- Python

- Apache Spark

- Spark Streaming

- Apache Kafka

- Apache Hudi

- Apache Iceberg

- Docker

- Amazon ECS

- Terraform

- OOPS

- Data Lake Architecture

- Real-Time Data Processing

- Distributed Data Processing

- Cloud Data Engineering

Roles & Responsibilities :

- Design, develop, and maintain scalable data engineering solutions on AWS.

- Build and optimize ETL/ELT pipelines using AWS Glue, Glue ETL, and PySpark.

- Develop high-performance real-time streaming applications using Spark Streaming and Apache Kafka.

- Design and implement scalable data lake solutions using Apache Hudi and Apache Iceberg.

- Process and analyze high-volume, high-velocity datasets using Amazon EMR.

- Develop reusable, efficient, and maintainable data processing frameworks.

- Create, optimize, and manage Glue Data Catalog for metadata management.

- Implement infrastructure automation using Terraform.

- Containerize applications using Docker and deploy workloads on Amazon ECS.

- Monitor, troubleshoot, and optimize data pipelines for performance, scalability, and reliability.

- Collaborate with data scientists, analysts, and cross-functional teams to deliver robust data platforms.

- Follow software engineering best practices including OOPS principles, code reviews, testing, and documentation.

- Ensure data quality, security, governance, and compliance across data platforms.

- Optimize Spark jobs and distributed workloads for maximum efficiency.

- Participate in architecture discussions and contribute to cloud modernization initiatives.

Required Qualifications :

- Bachelor's or Master's degree in Computer Science, Information Technology, Software Engineering, or a related field.

- 8+ years of experience in Data Engineering or Big Data development.

- Strong expertise in Python, PySpark, and Apache Spark.

- Hands-on experience with AWS Glue, Glue ETL, Glue Data Catalog, and Amazon EMR.

- Strong experience building real-time streaming solutions using Apache Kafka and Spark Streaming.

- Practical knowledge of Apache Hudi and Apache Iceberg.

- Experience with Docker containers and Amazon ECS.

- Strong understanding of Terraform for Infrastructure as Code (IaC).

- Solid understanding of distributed computing concepts and cloud-native architectures.

- Excellent problem-solving and debugging skills.

- Strong communication and collaboration abilities.

Preferred Skills :

- AWS Cloud certifications.

- Certifications in Apache Spark, Kafka, Docker, or Terraform.

- Experience designing enterprise-scale data lake architectures.

- Exposure to modern DevOps and CI/CD practices.

- Knowledge of performance tuning and optimization for Spark workloads.

- Experience working in Agile/Scrum environments.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...