Posted on: 05/06/2026
Role Overview :
We are building a next-generation data platform for a leading Media & Entertainment organization to unify Customer Experience (CX) data across multiple sources-including application reviews, social media platforms, and customer interaction channels.
This platform will ingest, process, and transform large-scale data into a modern data lake architecture, and power an intelligent insights layer using advanced analytics and Generative AI. The goal is to enable real-time, actionable insights that improve customer engagement and business decision-making.
Key Responsibilities :
- Design and build scalable data pipelines to ingest CX data from multiple sources such as app reviews, social media (Twitter, Instagram), and internal systems
- Implement robust ETL/ELT pipelines using Apache Spark and orchestrate workflows using Apache Airflow
- Develop and maintain data lake architecture following the Medallion (Bronze, Silver, Gold) pattern
- Work with Databricks on AWS to build and optimize distributed data processing systems
- Build an intelligent data layer leveraging analytics and Generative AI for actionable insights
- Design and implement RAG (Retrieval-Augmented Generation) pipelines for contextual insights
- Integrate LLM orchestration frameworks and Model Context Protocol (MCP) for scalable AI workflows
- Utilize AWS Bedrock (or equivalent platforms) to deploy and manage foundation models
- Ensure data quality, governance, security, and observability across pipelines
- Collaborate with data scientists, product teams, and business stakeholders to deliver high-impact solutions
Required Skills & Experience :
- Strong experience in Python & PySpark for data engineering (Must)
- Strong experience with Databricks platform for data processing (Must)
- Hands-on experience with Apache Spark for large-scale data processing
- Expertise in building and managing data pipelines using Apache Airflow
- Experience with AWS ecosystem (S3, Lambda, Glue, etc.)
- Solid understanding of data lake architectures and Medallion architecture
- Experience working with streaming and batch data pipelines
- Strong knowledge of REST/GraphQL APIs and data integration patterns
- Familiarity with data governance, cataloging, and monitoring tools
GenAI & Advanced Capabilities :
- Strong understanding of RAG (Retrieval-Augmented Generation) architectures
- Exposure to Model Context Protocol (MCP)
- Hands-on experience with AWS Bedrock
- Working knowledge of AWS Nova Pro
- Ability to design intelligent data products combining analytics + GenAI
Good to Have :
- Experience in Media & Entertainment or digital platforms
- Knowledge of social media analytics and sentiment analysis
- Exposure to real-time data processing frameworks (Kafka, Kinesis, etc.)
- Understanding of MLOps and model lifecycle management
Why Join Us :
- Opportunity to build a cutting-edge data + AI platform from the ground up
- Work on high-impact CX analytics influencing millions of users
- Exposure to modern data stack + Generative AI technologies
- Collaborative and innovation-driven environment
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
ML / DL Engineering
Job Code
1642134