Posted on: 04/06/2026
Job Description :
Role : Senior Big Data Engineer (Python & PySpark)
Role Summary :
We are looking for a Senior Big Data Engineer with a powerhouse combination of Python and PySpark expertise.
You will be responsible for designing, developing, and maintaining large-scale data processing systems. This role is ideal for a candidate who thrives in a complex data ecosystem and possesses a deep understanding of distributed computing and Object-Oriented Programming (OOPs).
Technical Pillars of the Role :
1. Core Python & Development :
- Expertise in Python development, including robust OOPs (Object-Oriented Programming) concepts.
- Advanced SQL skills for complex data querying and manipulation.
- Proficiency in Shell Scripting and Unix environments for automation and system-level tasks.
2. PySpark & Distributed Computing :
- Extensive experience with Spark and PySpark for large-scale data processing.
- Familiarity with Scala and traditional MapReduce frameworks.
- Hands-on experience with Hive for data warehousing and schema management.
3. Big Data Ecosystem & Orchestration :
- Deep knowledge of HDFS architecture and storage optimization.
- Proven experience using Sqoop for data ingestion and Flume for streaming data.
- Workflow orchestration using Oozie and enterprise scheduling via Autosys.
Key Responsibilities :
- Data Pipeline Development : Build and scale data pipelines using PySpark to process multi-terabyte datasets.
- Optimization : Tune Spark jobs for performance and resource efficiency within a Hadoop cluster.
- System Integration : Seamlessly move data between RDBMS and HDFS using Sqoop and Python scripts.
- Automation : Develop Unix Shell scripts to automate repetitive ETL tasks and monitor job health.
- Mentorship : Provide technical guidance to junior developers and contribute to code reviews and architecture design.
Candidate Requirements :
- Education : B.E./B.Tech/MCA in Computer Science or a related field.
- Experience : Minimum of 5 years in a Big Data environment, with at least 3 years focused on PySpark.
- Analytical Skills : Ability to troubleshoot complex distributed system issues and data bottlenecks.
- Communication : Strong interpersonal skills, essential for collaborating with cross-functional teams and stakeholders during the Hyderabad drive and beyond.
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Big Data / Data Warehousing / ETL
Job Code
1641888