Posted on: 19/05/2026
Description :
About the Role :
We are hiring experienced Python Engineers with strong expertise in Data Engineering, PySpark, and Spark for one of our esteemed clients. The ideal candidate will have hands-on experience building scalable data pipelines, processing large datasets, and optimizing distributed data workflows in enterprise environments.
Candidates who are currently serving notice period or can join within 30 days will be highly preferred.
Key Responsibilities :
- Design, develop, and maintain scalable data pipelines using Python and Apache Spark for large-scale data processing.
- Build and optimize ETL/ELT workflows to ingest, transform, and load data from structured and unstructured data sources.
- Work extensively with distributed data processing frameworks such as Spark (PySpark/Scala)
for handling big data workloads.
- Implement efficient data modeling strategies and optimize storage solutions across data lakes and data warehouses.
- Ensure data quality, governance, validation, and monitoring through automated checks and controls.
- Collaborate with Data Scientists, Analysts, and Business teams to deliver reliable datasets and actionable insights.
- Optimize Spark job performance using partitioning, caching, tuning, and query optimization techniques.
- Develop and manage workflow orchestration using tools such as Apache Airflow or equivalent schedulers.
- Work with SQL and NoSQL databases including PostgreSQL, MySQL, MongoDB, Hive, or similar technologies.
- Troubleshoot and resolve performance bottlenecks in data processing systems.
- Follow engineering best practices, coding standards, and agile development methodologies.
Required Skills & Qualifications :
- 5+ years of experience in Python Development/Data Engineering.
Strong hands-on experience in :
- Python
- PySpark
- Apache Spark
- SQL
- Experience building and maintaining ETL/ELT pipelines.
- Strong understanding of distributed computing and big data processing.
- Experience with workflow orchestration tools such as Airflow.
- Knowledge of data warehousing and data lake concepts.
- Hands-on experience with relational and NoSQL databases.
- Good understanding of performance tuning and optimization techniques in Spark.
- Strong analytical, debugging, and problem-solving skills.
- Excellent communication and collaboration abilities.
Preferred Skills :
- Exposure to cloud platforms such as AWS, Azure, or GCP.
- Experience with Scala is an added advantage.
- Familiarity with CI/CD pipelines and DevOps practices in data engineering environments.
Mandatory Requirements :
- Candidates should be available to join within 30 days or currently serving notice period.
- Final round of interview will be conducted Face-to-Face.
- Candidates must possess all employment and educational documents.
- PF should be deducted in current/previous organizations and UAN must be active and valid.
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1637047