Posted on: 23/06/2026
Senior Data Engineer Big Data & GenAI
Location: Chennai
Experience: 24 Years
Employment Type: Full-Time
Role Overview:
We are looking for a highly skilled and hands-on Data Engineer with strong expertise in Big Data technologies, distributed data processing, and backend engineering. The ideal candidate should have practical experience in building scalable batch and real-time data pipelines using frameworks such as PySpark, Kafka, and streaming platforms. The candidate should also possess strong Python programming skills and experience building APIs/microservices using modern backend frameworks. This role is best suited for engineers who enjoy solving large-scale data engineering problems and working closely with AI/GenAI-driven applications. We are specifically looking for candidates with strong hands-on coding and distributed systems experience, and not candidates whose experience is predominantly limited to traditional ETL tools, low-code cloud pipelines, or Snowflake-centric workflows.
Key Responsibilities:
- Design, develop, and optimize scalable batch and streaming data pipelines.
- Build large-scale distributed data processing systems using PySpark.
- Develop real-time streaming solutions using Kafka and/or Apache Flink.
- Build and maintain APIs and backend services using Python-based microservice frameworks.
- Work with structured and semi-structured datasets across multiple storage systems.
- Optimize pipeline performance, scalability, and fault tolerance.
- Collaborate with AI/ML and GenAI teams to support LLM and RAG-based applications.
- Participate in architecture discussions, debugging, code reviews, and production deployments.
- Ensure reliability, monitoring, and maintainability of data engineering systems.
Mandatory Skills:
- Strong hands-on experience with PySpark and distributed data processing (must-have).
- Strong programming expertise in Python.
- Hands-on experience building APIs and backend services using:
1. Flask
2. FastAPI
3. Or similar microservice frameworks
- Good understanding of Big Data ecosystem and streaming architectures.
- Experience with:
1. Apache Kafka
2. Apache Flink (preferred)
3. PostgreSQL
- Experience building end-to-end batch and real-time data engineering pipelines.
- Good understanding of scalable system design, partitioning, optimization, and distributed processing concepts.
- Understanding of microservices architecture and API integrations.
- Experience working in Linux-based environments.
- Strong analytical and problem-solving skills.
GenAI / AI Engineering Expectations:
- Candidates should also have practical exposure to modern GenAI concepts and tools, including:
1. Working knowledge of LLMs and prompt engineering.
2. Understanding of RAG (Retrieval-Augmented Generation) concepts.
3. Experience using GenAI frameworks/tools such as:
i. LangChain
ii. LlamaIndex (good to have)
- Exposure to vector databases, embeddings, or AI-assisted workflows is a plus.
Good to Have / Bonus Skills:
- Experience with QuestDB.
- Exposure to real-time analytics and streaming systems.
- Understanding of distributed systems internals.
- Experience with Docker/containerization and deployment workflows.
- Familiarity with monitoring and logging frameworks.
- Knowledge of cloud platforms is an added advantage.
Important Note:
This role is focused on core data engineering, distributed systems, and backend engineering. Candidates with predominantly:
1. Snowflake-only experience
2. Low-code ETL pipeline development
3. Pure cloud ETL orchestration backgrounds
may not be the right fit for this requirement unless they also possess strong hands-on Big Data engineering and programming experience.
Preferred Candidate Profile:
- Strong ownership mindset and problem-solving ability.
- Comfortable working independently on complex engineering problems.
- Ability to work in fast-paced product/research-oriented environments.
- Passion for learning modern AI and data technologies.
- Good communication and collaboration skills.
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1647450