HamburgerMenu
hirist

PySpark Developer

SP Staffing Services
5 - 10 Years
Multiple Locations

Posted on: 02/09/2026

Job Description

Job Description :


Key Responsibilities :


- Design and develop scalable data processing pipelines using PySpark.

- Develop and optimize ETL/ELT workflows for large and complex datasets.

- Write efficient PySpark and Python code for data extraction, transformation, and loading.

- Perform data cleansing, transformation, aggregation, and validation using PySpark.

- Develop complex SQL queries and optimize data processing workloads.

- Work with structured and unstructured data from databases, APIs, files, and streaming sources.

- Optimize Spark jobs for performance, scalability, memory utilization, and cost efficiency.

- Implement data quality checks, validation, monitoring, and error-handling mechanisms.

- Troubleshoot and resolve data pipeline and production issues.

- Collaborate with Data Engineers, Data Architects, Data Scientists, Analysts, and business stakeholders.

- Participate in code reviews, testing, deployment, and CI/CD activities.

- Ensure data pipelines follow security, governance, and compliance standards.

- Contribute to the design and implementation of reliable, maintainable, and production-ready data solutions.

Required Skills :

- 5 - 10 years of experience in Data Engineering / Big Data development.

- Strong hands-on experience with PySpark and Apache Spark.

- Strong proficiency in Python.

- Strong knowledge of SQL and query optimization.

- Experience with ETL/ELT pipeline development.

- Good understanding of distributed computing and data processing concepts.

- Experience working with large-scale datasets.

- Knowledge of data warehousing, data lakes, and data modelling.

- Experience with cloud platforms such as AWS, Azure, or GCP is preferred.

- Strong debugging, troubleshooting, and performance-tuning skills.

Preferred Skills :

- Experience with Spark SQL, DataFrames, RDDs, and Spark optimization.

- Knowledge of Kafka or other streaming technologies.

- Experience with data platforms such as Databricks, Snowflake, or Hadoop.

- Exposure to Airflow or other workflow orchestration tools.

- Experience with Git, CI/CD, and DevOps practices.

- Understanding of data security, governance, and data quality frameworks.

Preferred Candidate Profile :

- Strong analytical and problem-solving abilities.

- Experience building production-grade and scalable data pipelines.

- Ability to work independently and collaborate effectively with cross-functional teams.

- Strong communication and stakeholder-management skills.

- Good understanding of modern data engineering practices and cloud technologies.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...