Posted on: 17/09/2026
Role & Responsibilities :
- Design, develop, and maintain scalable and high-performance big data processing pipelines.
- Develop and optimize data pipelines using technologies such as Apache Spark, PySpark, Kafka, Hadoop, Hive, and Airflow.
- Build robust ETL/ELT workflows to process large volumes of structured and unstructured data.
- Work with cloud-based data platforms such as AWS, Azure, or GCP and implement cloud-native data solutions.
- Design and implement data lakes, data warehouses, and distributed data processing architectures.
- Optimize Spark jobs, SQL queries, data pipelines, and storage for performance, scalability, and cost efficiency.
- Develop real-time and batch data processing solutions using Kafka and Spark Streaming/Structured Streaming.
- Ensure data quality, accuracy, availability, security, and compliance across data pipelines.
- Troubleshoot production issues, perform root-cause analysis, and implement permanent fixes.
- Collaborate with Data Scientists, Data Analysts, Software Engineers, Architects, and Business teams to understand data requirements and deliver scalable solutions.
- Participate in system design, architecture discussions, code reviews, and technical documentation.
- Establish and follow engineering best practices for version control, CI/CD, testing, monitoring, and deployment.
- Mentor junior engineers and provide technical guidance to the team.
- Stay current with emerging big data, cloud, data engineering, and distributed computing technologies.
Preferred Candidate Profile :
- 5+ years of experience in Big Data Engineering, Data Engineering, or a closely related role.
- Strong hands-on experience with Apache Spark / PySpark and distributed data processing.
- Strong programming skills in Python, Scala, or Java.
- Strong understanding of data lake, data warehouse, dimensional modeling, and distributed system concepts.
- Hands-on experience with at least one major cloud platform: AWS, Microsoft Azure, or Google Cloud Platform.
- Experience with modern data platforms such as Databricks, Snowflake, Delta Lake, or similar technologies is preferred.
- Knowledge of Docker, Kubernetes, Git, CI/CD, and DevOps practices is an advantage.
- Strong understanding of data quality, data governance, security, monitoring, and performance optimization.
- Ability to analyze complex technical problems and develop scalable, reliable, and maintainable solutions.
- Good communication and collaboration skills, with the ability to work effectively across technical and business teams.
- Experience mentoring engineers or taking ownership of technical design and delivery is preferred.
- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field is preferred.
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1672366