HamburgerMenu
hirist

Infosys - Senior Data Engineer - Python/Big Data

Infosys Limited
8 - 20 Years
Multiple Locations

Posted on: 25/08/2026

showcase-imageshowcase-image

Job Description

Role Overview:

We are looking for a highly skilled Data Engineer with expertise in designing, developing, and optimizing enterprise-scale data platforms and pipelines. The ideal candidate should possess strong experience across data engineering, cloud-native architectures, modern data platforms, streaming technologies, and AI-ready data ecosystems.

Candidates with exposure to AI, Generative AI, DataOps, MLOps, Lakehouse architectures, and Real-Time Analytics will be highly preferred.

Key Responsibilities:

Data Platform Engineering:

- Design, develop, and maintain scalable data pipelines for batch and real-time processing.

- Build enterprise-grade data platforms supporting analytics, AI, and business intelligence workloads.

- Develop reusable ingestion, transformation, and orchestration frameworks.

- Implement scalable ETL/ELT solutions across structured and unstructured data sources.

Data Architecture & Modeling:

- Design modern Lakehouse and Data Mesh architectures.

- Build optimized data models supporting reporting, analytics, and machine learning.

- Establish enterprise data standards, governance, and lineage frameworks.

- Ensure data quality, consistency, and regulatory compliance.

Cloud Data Engineering:

- Build cloud-native data platforms on AWS, Azure, or GCP.

- Design scalable solutions leveraging managed data services.

- Optimize performance, reliability, and cost efficiency of cloud data workloads.

- Implement Infrastructure as Code and automated deployment pipelines.

Real-Time Data Engineering:

- Build event-driven architectures and streaming data solutions.

- Develop real-time ingestion and processing pipelines.

- Implement high-throughput and low-latency data processing systems.

AI & GenAI Data Foundation:

- Build AI-ready data ecosystems supporting machine learning and Generative AI workloads.

- Engineer pipelines for feature stores, vector databases, and RAG architectures.

- Enable enterprise-scale data preparation for AI/ML models.

- Collaborate with AI Engineers and Data Scientists for model operationalization.

Engineering Excellence:

- Drive best practices for DataOps, DevOps, CI/CD, testing, and observability.

- Participate in architecture reviews and technical design discussions.

- Contribute to reusable assets, accelerators, and intellectual property development.

- Mentor junior engineers and support technical leadership initiatives.

Required Technical Skills:

Data Engineering:

- Python, PySpark, SQL (Advanced), Data Modeling, ETL/ELT Development, Data Warehousing Concepts

Big Data Technologies:

- Apache Spark, Databricks, Hadoop Ecosystem, Delta Lake, Apache Iceberg, Apache Hudi

Data Integration & Streaming:

- Apache Kafka, Apache Flink, Apache Airflow, Azure Data Factory, AWS Glue, Stream Processing Frameworks

Cloud Platforms:

- Azure (Data Factory, Synapse, Data Lake, Fabric, Databricks), AWS (S3, Glue, EMR, Redshift, Kinesis, Lambda), GCP (BigQuery, Dataflow, Dataproc, Pub/Sub)

Databases:

- PostgreSQL, SQL Server, Oracle, MongoDB, Cassandra, Snowflake

DataOps & DevOps:

- Git, CI/CD Pipelines, Jenkins, GitHub Actions, Terraform, Docker, Kubernetes

Preferred Skills:

- Generative AI, LLMs, RAG, Vector Databases (Pinecone, Milvus, Weaviate), MLOps, Data Governance, MDM, Data Quality Frameworks, Knowledge Graphs, Agentic AI Data Platforms

Educational Qualification:

- Bachelor's or Master's degree in Computer Science, Engineering, Data Science, Information Technology, or related disciplines.

Candidate Profile:

- Strong coding and problem-solving skills.

- Deep understanding of distributed data systems.

- Passionate about building large-scale data platforms.

- Ability to work independently and thrive in ambiguous environments.

- Experience working in Agile and cloud-native engineering teams.

- Demonstrate innovation, ownership, and engineering excellence.

Success Metrics:

- Scalable and reliable data platform delivery.

- Data quality and governance compliance.

- Performance optimization and cost efficiency improvements.

- Contribution to reusable frameworks, accelerators, and IP.

- Enablement of Analytics, AI, and GenAI use cases.

- Technical leadership and mentorship impact.

The job is for:

Women candidates preferred
info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...