Posted on: 23/03/2026
GCP Dataproc Specialist
Role Overview :
As a GCP Dataproc Specialist, you will be instrumental in designing, developing, and maintaining robust data processing pipelines using Google Cloud Platform's Dataproc service. You will collaborate closely with data scientists, data engineers, and business stakeholders to understand their data processing needs and translate them into efficient and scalable solutions. Your work will directly impact the organization's ability to derive valuable insights from large datasets, enabling data-driven decision-making and improved business outcomes.
Key Responsibilities :
- Design and implement scalable and efficient data processing pipelines using GCP Dataproc, Spark, and Hadoop for various data ingestion, transformation, and analysis requirements.
- Develop and maintain Python, SQL, and Hive scripts to process and analyze large datasets stored in Google Cloud Storage and other data sources, ensuring data quality and accuracy.
- Optimize Spark applications for performance and cost-effectiveness, leveraging best practices for resource allocation, data partitioning, and caching to meet stringent SLAs.
- Collaborate with data engineers to build and maintain data infrastructure, including data lakes, data warehouses, and data pipelines, ensuring data availability and reliability.
- Troubleshoot and resolve issues related to Dataproc clusters, Spark applications, and data pipelines, providing timely support to data scientists and other stakeholders.
- Implement and maintain data security and governance policies, ensuring compliance with industry regulations and company standards.
Required Skillset :
- Demonstrated expertise in designing, developing, and deploying data processing solutions using GCP Dataproc, Hadoop, and Spark.
- Proven ability to write efficient and scalable Python, SQL, and Hive scripts for data manipulation and analysis.
- Strong understanding of data warehousing concepts, data modeling techniques, and ETL processes.
- Excellent problem-solving skills and the ability to troubleshoot complex issues in a distributed computing environment.
- Effective communication and collaboration skills, with the ability to work effectively with cross-functional teams.
- Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
- Experience with Apache Spark is a must.
- Adaptable to work in a remote/hybrid environment.
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1622646