Posted on: 25/08/2026
About the Role :
We are looking for a Senior Data Engineer / AI Data Platform Engineer with strong expertise in modern data engineering and AI data platforms.
The ideal candidate should have hands-on experience building scalable batch and streaming data pipelines using Python/Scala, Apache Spark, Apache Airflow, and Databricks on AWS, along with strong exposure to RAG architectures and LLM orchestration frameworks.
The role will involve designing modern data lake architectures, integrating enterprise and AI data sources, and enabling GenAI use cases through technologies such as AWS Bedrock, AWS Nova Pro, MCP, and RAG.
Key Responsibilities :
Data Engineering & Pipeline Development :
- Design, develop, and maintain scalable data pipelines using Python and/or Scala.
- Build and optimize data processing solutions using Apache Spark.
- Develop and manage workflow orchestration using Apache Airflow.
- Build and manage data pipelines on Databricks running on AWS.
- Develop both batch and streaming data pipelines.
- Implement reliable and scalable data ingestion and transformation processes.
- Work with AWS services including Amazon S3, AWS Lambda, and AWS Glue.
Data Lake & Data Platform Architecture :
- Design and implement modern data lake architectures.
- Work with the Medallion Architecture across : 1. Bronze, 2. Silver, 3. Gold.
- Develop scalable data models and transformation frameworks.
- Ensure data platforms are reliable, maintainable, and optimized for downstream analytics and AI applications.
- Implement appropriate data governance, cataloging, and monitoring practices.
AI Data Platform & GenAI :
- Build data platforms that support GenAI and LLM-based applications.
- Design and implement RAG (Retrieval-Augmented Generation) architectures.
- Work with LLM orchestration frameworks to build AI-enabled data workflows.
- Integrate AI applications with enterprise data sources.
- Work with Model Context Protocol (MCP).
- Develop and integrate solutions using AWS Bedrock.
- Work with AWS Nova Pro for AI/LLM-driven use cases.
- Support data preparation, retrieval, orchestration, and integration requirements for LLM applications.
APIs & Data Integration :
- Design and implement data integration patterns using REST and GraphQL APIs.
- Integrate data from multiple internal and external systems.
- Develop reliable data ingestion and transformation workflows.
- Ensure appropriate handling of data quality, availability, and integration requirements.
Streaming & Real-Time Data :
- Design and develop real-time data processing solutions.
- Work with streaming technologies such as : 1. Apache Kafka, 2. Amazon Kinesis.
- Support real-time analytics and AI/ML use cases.
Data Governance & Monitoring :
- Implement data governance and data quality practices.
- Work with data cataloging and monitoring tools.
- Ensure data pipelines are observable, reliable, and maintainable.
- Monitor pipeline performance and troubleshoot data processing issues.
Mandatory Skills :
- Strong hands-on experience with Python and/or Scala.
- Strong experience with Apache Spark.
- Strong experience with Apache Airflow.
- Hands-on experience with Databricks on AWS.
- Strong experience with RAG architectures.
- Experience with LLM orchestration frameworks.
- Strong understanding of the AWS ecosystem, particularly S3, Lambda, and Glue.
- Experience with Data Lake architecture.
- Strong understanding of Medallion Architecture Bronze, Silver, Gold.
- Experience building batch and streaming data pipelines.
- Experience with REST/GraphQL APIs and data integration patterns.
- Understanding of data governance, cataloging, and monitoring.
- Experience with Model Context Protocol (MCP).
- Experience with AWS Bedrock.
- Experience with AWS Nova Pro.
Preferred Skills :
- Experience in Media & Entertainment or digital platform environments.
- Experience with social media analytics.
- Experience with sentiment analysis.
- Hands-on experience with Kafka or Kinesis for real-time processing.
- Experience with MLOps and model lifecycle management.
Ideal Candidate :
- The ideal candidate should combine strong data engineering fundamentals with modern AI/GenAI platform expertise.
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1665830