Posted on: 08/10/2026
Core Responsibilities :
- Design and implement large-scale distributed data platforms and data lakes.
- Manage data ingestion from diverse sources and file formats.
- Perform Big Data orchestration using Airflow, Spark on Kubernetes, Yarn, and Oozie.
- Optimize SQL-Hive and Spark queries for performance.
- Ensure data validation and quality.
- Support 24x7 operations as per the rota.
Technical Requirements :
- Advanced Scala (Functional Programming, Case classes, Complex Data Structures & Algorithms).
- Big Data Unit, System, Integration & Regression Testing.
- DevOps : Jenkins, Maven/sbt, Git, GitHub Actions.
- Agile processes & tools : Jira & Confluence.
Added Advantage :
- Data Streaming : Kafka and Spark Streaming.
- Data Lake & Medallion Architecture.
- Shell Scripting & Automation (Ansible).
- Cloud Native Storage (Ceph or S3) and Cloud Data Engineering (Azure, AWS & GCP).
- SQL & NoSQL DB : Hive, Cassandra, Dremio, Presto.
- Infrastructure : Kubernetes, Docker and related container technologies.
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1677366