HamburgerMenu
hirist

Data Engineer - ETL/PySpark

HR Works Consultancy
6 - 10 Years
Pune

Posted on: 16/07/2026

Job Description

Data Pipeline Development & Operations :

- Design, build, and operate scalable and reliable data pipelines on the Databricks platform

- Develop end-to-end data workflows from ingestion through transformation to consumption

- Implement robust error handling, monitoring, and alerting mechanisms

- Ensure data pipeline reliability, performance, and maintainability

- Optimize pipeline performance through efficient Spark job design and cluster configuration

- Manage and orchestrate complex data workflows using Databricks Jobs and workflows

Legacy Code Modernization :

- Refactor legacy code and data pipelines to PySpark for improved performance and scalability

- Migrate traditional ETL processes to modern ELT patterns on Databricks

- Assess existing codebases and identify opportunities for optimization and modernization

- Ensure backward compatibility and data integrity during migration processes

- Document refactoring approaches and create migration playbooks

- Collaborate with stakeholders to minimize disruption during code transitions

Data Engineering Excellence :

- Implement data quality checks and validation frameworks

- Design and maintain Delta Lake tables with appropriate optimization strategies

- Develop reusable code libraries and frameworks for common data engineering tasks

- Follow software engineering best practices including version control, testing, and CI/CD

- Participate in code reviews and provide constructive feedback to team members

- Troubleshoot and resolve data pipeline issues in production environments

Collaboration & Knowledge Sharing :

- Work closely with data architects, analysts, and business stakeholders

- Collaborate with Infrastructure (Infra), Applications (Apps), and Cyber teams

- Share knowledge and best practices with Team NCS

- Mentor junior data engineers on PySpark and Databricks technologies

- Document technical solutions and maintain comprehensive documentation

Essential Technical Skills :

- Data Engineering: Strong foundation in data engineering principles, ETL/ELT processes, and data pipeline design patterns

- PySpark: Proven hands-on experience developing data pipelines using PySpark, including DataFrames API, Spark SQL, and performance optimization

- Databricks Platform: Practical experience with Databricks workspace, cluster management, notebooks, and job orchestration

- Workspace AI Agent: Knowledge of Databricks Workspace AI Agent capabilities and integration

- Data Modelling: Experience implementing data models including dimensional modeling, data vault, or lakehouse architectures

- Delta Lake: Understanding of Delta Lake features including ACID transactions, schema evolution, and optimization techniques

- Python: Strong Python programming skills for data processing and automation 5+ years of relevant experience

- Strong foundation in data engineering principles, ETL/ELT processes, and data pipeline design patterns

- Proven hands-on experience developing data pipelines using PySpark, including DataFrames API, Spark SQL, and performance optimization

- Min 2 to 3 yrs exp in Databricks Platform: Practical experience with Databricks workspace, cluster management, notebooks, and job orchestration

- Experience implementing data models including dimensional modeling, data vault, or lakehouse architectures

- Understanding of Delta Lake features including ACID transactions, schema evolution, and optimization techniques

- Strong Python programming skills for data processing and automation

- Experience with cloud platforms (Azure, AWS, or GCP) mandatory to have at least one certification

- Databricks Certified Data Engineer Associate OR Databricks Certified Data Engineer Professional

Notice Period - Immediate to 30 days

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...