HamburgerMenu
hirist

Apache Spark & Scala Developer

SysMind
6 - 11 Years
Anywhere in India/Multiple Locations

Posted on: 09/06/2026

Job Description

Role Overview:

This pivotal role involves architecting, developing, and optimizing large-scale data processing solutions that power critical business intelligence and analytical capabilities.


You will be instrumental in designing robust, scalable, and high-performance data pipelines using Apache Spark and Scala, tackling complex data challenges across diverse datasets.


Working closely with data scientists, product managers, and fellow engineers, you will transform raw data into actionable insights, directly impacting strategic decision-making and enhancing customer experiences.


Your contributions will be central to building and maintaining a cutting-edge data platform that drives innovation and business growth.

Key Responsibilities:

- Design and implement highly scalable and efficient data processing pipelines using Apache Spark and Scala for both batch and real-time data, ensuring data quality and availability for downstream applications.

- Develop and optimize complex SQL queries and data models within the Hadoop ecosystem (HDFS, YARN) to support robust data storage, retrieval, and transformation, meeting stringent performance requirements.

- Collaborate with cross-functional teams, including data scientists and business analysts, to translate complex data requirements into technical specifications and deliver impactful data solutions.

- Architect and manage data lake components, leveraging HDFS and other distributed storage technologies, to ensure data integrity, security, and accessibility across the enterprise.

- Lead performance tuning, troubleshooting, and debugging efforts for Spark applications and Hadoop jobs, identifying bottlenecks and implementing strategic optimizations to enhance system efficiency and reliability.

- Contribute to the evolution of our data architecture, promoting best practices in data engineering, and mentoring junior team members to foster a culture of technical excellence.

Required Skillset:

- Deep expertise in Apache Spark for large-scale data processing, including Spark SQL, Spark Streaming, and advanced performance tuning techniques.

- Exceptional proficiency in Scala programming for developing robust, maintainable, and high-performance data applications.

- Strong command over SQL for complex data manipulation, querying, and optimization across various data sources.

- Extensive hands-on experience with the Hadoop ecosystem, including HDFS, YARN, and related distributed computing technologies.

- Demonstrated ability in data modeling principles and practices, capable of designing scalable and efficient data structures for analytical and operational needs.

- Proven problem-solving capabilities, with a proactive approach to identifying and resolving complex technical challenges in a distributed data environment.

- Excellent communication and interpersonal skills, adept at articulating technical concepts clearly to both technical and non-technical stakeholders.

- A Bachelor's or Master's degree in Computer Science, Engineering, or a related quantitative field from a premier institution.

- Adaptability to work effectively in a dynamic, fast-paced environment, managing multiple priorities and evolving project requirements.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...