Posted on: 22/07/2026
Pyspark Data Engineer
Requirements :
- Hands-on expertise in designing, building, and maintaining Apache Spark pipelines in production environments.
- Proven experience building and scaling data ingestion frameworks that integrate data from multiple source systems, with a focus on reliability, reusability, and scalability.
- Deep understanding of Spark architecture (driver/executors, DAG, partitioning, shuffles, caching, cluster resource management) and experience operating pipelines at scale, including data transformations on datasets ~500 GB+.
- Strong understanding of Oracle SQL and HDFS, including handling file formats and applying appropriate data cleansing, normalization, and formatting to produce curated output datasets.
- Ability to write Python, Pyspark, and shell scripts to process, transform, and automate data workflows. The candidate should be good in writing application programs and automation manual data processing steps using python.
Skills :
- Strong hands-on experience in PySpark, Python, and SQL.
- Experience designing and optimizing Spark-based ETL/ELT pipelines and data processing jobs.
- Strong understanding to a BigQuery.
- Strong understanding of data quality, governance, observability, and performance tuning.
- Good collaboration, debugging, and Agile delivery skills.
Experience :
- Bachelors or Masters degree plus 6+ years of data engineering experience with strong PySpark expertise.
Role Overview :
We are seeking a highly skilled Senior Data Engineer to join our dynamic data platform team in Bangalore, Hyderabad, Chennai, or Pune. In this role, you will architect, build, and maintain scalable data pipelines that process massive datasets to drive critical business intelligence and machine learning initiatives.
You will collaborate closely with cross-functional teams, including Data Scientists, Product Managers, and Business Analysts, to translate complex business requirements into robust technical solutions. By optimizing data architecture and ensuring high-quality data availability, you will directly influence strategic decision-making and enhance the performance of our data-driven products for global stakeholders.
Key Responsibilities :
- Design and implement high-performance data processing pipelines using PySpark to handle large-scale batch and real-time data ingestion, ensuring seamless integration across distributed systems.
- Develop and optimize complex Oracle SQL queries and stored procedures to support high-volume data retrieval, ensuring maximum efficiency and minimal latency for end-user applications.
- Automate data workflows and build reusable data frameworks using Python to improve operational efficiency and reduce manual intervention for the engineering team.
- Collaborate with stakeholders to define data modeling strategies that align with business goals, ensuring data integrity and consistency across the enterprise ecosystem.
- Conduct rigorous code reviews and performance tuning of existing data processes to identify bottlenecks and implement scalable improvements that support long-term business growth.
Required Skillset :
- Demonstrated expertise in building and managing large-scale data solutions using PySpark, with a deep understanding of distributed computing principles and performance optimization techniques.
- Advanced proficiency in Oracle SQL, including complex query optimization, database schema design, and performance tuning for high-concurrency environments.
- Strong programming capabilities in Python, with a focus on writing clean, maintainable, and production-ready code for data engineering tasks.
- Proven ability to communicate complex technical concepts to non-technical stakeholders, fostering a collaborative environment that bridges the gap between business needs and engineering execution.
- A proactive mindset with the ability to thrive in a hybrid work environment, demonstrating self-management and the capacity to adapt to evolving project priorities.
- A Bachelors or Masters degree in Computer Science, Information Technology, or a related quantitative field, supported by 6 - 8 years of hands-on experience in data engineering roles.
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1656370