Posted on: 17/07/2026
Description :
Desired Competencies (Technical/Behavioral Competency):
Must-Have :
- Build, maintain, and troubleshoot pipelines using GCP services (e.g., DataProc, Pub/Sub, Cloud Functions, Cloud Composer/Apache Airflow) to ingest, transform, and load data from various sources (relational databases, APIs, streaming data, flat files).
- Implement batch data processing solutions.
- Identify and resolve data-related issues, including data quality problems, pipeline failures, and performance bottlenecks.
- Sound programming knowledge on PySpark & SQL in terms of processing large amount of semi structured & unstructured data
- Working Knowledge on working with Avro, Parquet format files
- Knowledge on working on Hadoop Big Data platform and ecosystem
Good-to-Have :
- Knowledge on Jira, Agile, Sonar, Team city & CICD
- Any exposure / experience for an international Banking client / multi-vendor / multi geography teams
Responsibility of / Expectations from the Role :
- Design, implement, and optimize scalable, reliable, and secure data architectures on GCP, including data lakes, data warehouses, and streaming solutions.
- Develop and maintain data models for optimal storage, retrieval, and analysis.
- Build, maintain, and troubleshoot robust ETL/ELT pipelines using GCP services (e.g., Dataflow, Pub/Sub, Cloud Functions, Cloud Composer/Apache Airflow) to ingest, transform, and load data from various sources (relational databases, APIs, streaming data, flat files).
- Ensure data quality, consistency, and integrity by implementing validation frameworks, reconciliation checks, and monitoring dashboards.
- Perform performance tuning and cost optimization for GCP data pipelines, including Dataproc clusters, Dataflow jobs, and storage layers.
- Implement data security controls such as IAM policies, encryption (at rest and in transit), and access auditing in compliance with enterprise standards.
- Design fault-tolerant and highly available data pipelines with proper retry, alerting, and failure-handling mechanisms.
- Manage schema evolution and metadata using tools like Hive Metastore, BigQuery schema management, or Data Catalog.
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1655334