HamburgerMenu
hirist

Infosys - AWS Data Engineer

Infosys Limited
2 - 5 Years
Multiple Locations

Posted on: 26/08/2026

showcase-imageshowcase-image

Job Description

Role Overview :

The ideal candidate should have strong hands-on experience with AWS data services, PySpark, SQL, ETL/ELT development, and large-scale data processing. You will work closely with data analysts, data scientists, technology teams, and business stakeholders to deliver robust data solutions that support analytics, reporting, and business intelligence initiatives.

Key Responsibilities :

- Design, develop, and maintain scalable batch and real-time data pipelines using AWS cloud services.

- Build robust ETL/ELT workflows for data ingestion, transformation, validation, and loading from multiple structured and unstructured data sources.

- Develop and maintain cloud-based data lakes using Amazon S3 and related AWS services.

- Design and optimize data warehouse solutions using Amazon Redshift.

- Develop data integration and orchestration frameworks using AWS Glue, Lambda, EMR, and other AWS services.

- Develop PySpark-based data processing applications for large-scale distributed data workloads.

- Write optimized SQL queries for data transformation, analysis, and reporting requirements.

- Implement data ingestion frameworks to process data from databases, APIs, files, applications, and other enterprise systems.

- Optimize data pipelines for performance, scalability, reliability, and cost efficiency.

- Monitor data pipelines and troubleshoot failures, latency, data quality issues, and performance bottlenecks.

- Implement data validation and quality checks to ensure accuracy, consistency, completeness, and reliability of enterprise data.

- Build and maintain data processing solutions using AWS Athena for querying data stored in data lakes.

- Work with Amazon EMR for distributed processing and large-scale data transformation workloads.

- Implement appropriate data security, access controls, encryption, and governance practices across AWS data platforms.

- Automate infrastructure provisioning and deployment using Infrastructure as Code practices.

- Contribute to CI/CD processes for data engineering workflows and production deployments.

- Ensure data platforms are highly available, scalable, secure, and resilient.

- Collaborate with data analysts, data scientists, architects, application teams, and business stakeholders to understand data requirements.

- Participate in technical design discussions, code reviews, solution development, and production support.

- Maintain technical documentation related to data pipelines, data models, integrations, and platform architecture.

Technical Skills :

- Strong hands-on experience in AWS Data Engineering and cloud-native data solutions.

- Proficiency in AWS Glue, Amazon S3, Amazon Redshift, Amazon Athena, Amazon EMR, and AWS Lambda.

- Strong programming experience in Python, with hands-on expertise in PySpark.

- Strong SQL skills, including complex queries, joins, aggregations, window functions, and query optimization.

- Experience designing and implementing data lakes and enterprise data warehouse solutions.

- Strong understanding of ETL/ELT architecture, data ingestion, transformation, orchestration, and processing.

- Experience working with distributed data processing frameworks, particularly Apache Spark and PySpark.

- Understanding of data partitioning, file formats, schema management, and performance optimization.

- Experience with data pipeline monitoring, logging, troubleshooting, and failure recovery.

- Knowledge of cloud data security, access management, encryption, and data governance practices.

- Experience with Infrastructure as Code and automated deployment practices.

- Familiarity with CI/CD processes and Agile/Scrum development methodologies.

Preferred Experience :

- Experience working on enterprise-scale data platforms and high-volume data processing environments.

- Experience handling both batch and streaming data workloads.

- Exposure to modern data platform architecture and cloud-native data engineering practices.

- Experience optimizing AWS infrastructure and data workloads for performance and cost.

- Ability to independently troubleshoot complex data pipeline and production issues.

- Strong understanding of data engineering best practices, design patterns, and scalable architecture.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...