HamburgerMenu
hirist

Ekfrazo Technologies - Principal Data Engineer - Big Data & Spark Streaming

EKFRAZO TECHNOLOGIES PRIVATE LIMITED
7 - 10 Years
Multiple Locations

Posted on: 08/08/2026

Job Description

About the Role :

We are looking for an experienced Principal Data Engineer to design, build, and optimize scalable big data platforms and real-time data processing pipelines. The ideal candidate should have strong expertise in Data Modelling, Advanced SQL, Spark Streaming, Python, and Big Data ecosystems, with hands-on experience handling terabyte-scale datasets and building high-performance batch and streaming solutions.

Key Responsibilities :

- Design, develop, and optimize scalable data pipelines for batch and real-time data processing.

- Build and maintain robust Big Data platforms capable of handling terabyte-scale datasets.

- Develop Spark applications and optimize Spark SQL jobs for high-performance data processing.

- Design and implement semantic data models to support analytics and business intelligence.

- Process streaming data using Apache Spark Streaming and Kafka.

- Write and optimize complex SQL and Hive queries involving joins, UDFs, views, partitions, and large datasets.

- Build efficient data ingestion frameworks for structured, semi-structured, and unstructured data.

- Configure, schedule, and monitor workflows using Airflow and/or Oozie.

- Work with multiple file formats including ORC, AVRO, and Parquet.

- Develop cloud-based data solutions using AWS, Azure, or GCP.

- Collaborate with cross-functional teams to design scalable data architectures and analytics solutions.

- Troubleshoot performance bottlenecks and optimize distributed data processing workloads.

- Participate in code reviews and ensure adherence to engineering best practices.

Required Skills & Experience :

- 7+ years of experience in Data Engineering or Big Data Engineering.

- Strong expertise in :

1. Data Modelling

2. Advanced SQL

3. Semantic Modelling

4. Handling Terabyte-scale datasets

- Hands-on experience with :

1. Apache Spark

2. Spark SQL

3. Spark Streaming (Mandatory)

4. Python (Preferred) or Scala

5. Apache Kafka or other messaging platforms

6. Hive

7. Airflow and/or Oozie

- Strong knowledge of :

1. Batch and Real-time Streaming Data Processing

2. Big Data Ecosystems

3. Data Ingestion Frameworks

4. Distributed Computing

- Experience working with :

1. ORC

2. AVRO

3. Parquet

4. Unstructured Data

- Experience with cloud platforms such as AWS, Azure, or GCP.

- Experience with at least one modern data warehouse :

1. Snowflake

2. AWS Redshift

3. Google BigQuery

- Experience with NoSQL storage solutions such as Amazon S3 or similar object storage.

- Excellent analytical, debugging, and performance tuning skills.

Good to Have :

- AWS services such as EMR, S3, Redshift, ECS/EKS.

- GCP services such as Dataproc and Google Cloud Storage.

- Apache Iceberg.

- Hadoop MapReduce.

- Apache Flink.

- Kubernetes.

- ELK Stack (especially Elasticsearch).

- Experience working with large-scale Big Data clusters containing millions of records.

Interview Process :

- Round 1 : Technical Interview GlobalLogic Engineering Team

- Round 2 : Client Technical Interview

- Round 3 : Client Technical/Managerial Interview (Mandatory)

Why Join Us :

- Work on enterprise-scale Big Data and real-time streaming platforms.

- Build high-performance data solutions for global clients.

- Gain exposure to cloud-native data engineering and modern analytics technologies.

- Collaborate with experienced engineering teams on cutting-edge data transformation initiatives.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...