Posted on: 15/08/2026
About the Role :
We are seeking a highly accomplished Principal Engineer to lead the design and development of next-generation AI-ready Data Platforms, Real-Time Streaming Architectures, and Large-Scale Distributed Systems. The ideal candidate will possess deep expertise across Data Engineering, Kafka Streaming, Java-based backend systems, cloud-native architectures, and advanced data platforms that power analytics, machine learning, and Generative AI solutions.
As a Principal Engineer, you will drive technical strategy, influence architectural decisions, mentor engineering teams, and collaborate with product, platform, and AI teams to build scalable and resilient systems capable of processing petabyte-scale data workloads.
Key Responsibilities :
1. Technical Leadership :
- Define and drive enterprise-wide architecture for modern data platforms and AI ecosystems.
- Lead the design of highly scalable, fault-tolerant, and distributed data processing systems.
- Establish engineering standards, best practices, and reusable frameworks across teams.
- Influence technology roadmaps and strategic engineering investments.
2. Data Engineering & Data Platforms :
- Architect and implement large-scale batch and real-time data pipelines.
- Design modern Data Lake, Lakehouse, Data Warehouse, and Data Mesh architectures.
- Build robust data ingestion, transformation, governance, lineage, and observability frameworks.
- Enable AI/ML and GenAI workloads through scalable, high-performance data platforms.
3. Real-Time Streaming & Event-Driven Systems :
- Design enterprise-scale event streaming platforms using Apache Kafka.
- Build real-time analytics and event-driven architectures supporting millions of transactions.
- Develop streaming data pipelines and event-processing applications.
- Optimize throughput, latency, fault tolerance, and scalability of streaming ecosystems.
4. Distributed Systems Engineering :
- Design resilient microservices and distributed architectures.
- Build high-performance systems handling large-scale concurrent workloads.
- Solve challenges related to data consistency, reliability, scalability, and system performance.
- Drive observability, reliability engineering, and platform resiliency initiatives.
5. Artificial Intelligence & Data Ecosystem :
- Collaborate with AI and Data Science teams to operationalize AI solutions.
- Build data foundations for machine learning, LLM, Agentic AI, and Generative AI applications.
- Enable vector databases, feature stores, model-serving platforms, and AI data pipelines.
- Drive integration of AI capabilities into enterprise platforms and products.
6. Technical Mentorship :
- Mentor senior engineers and architects across multiple programs.
- Conduct architecture reviews and technical design assessments.
- Champion innovation and engineering excellence across the organization.
Required Skills & Experience :
1. Core Technologies :
- 12+ years of software engineering experience with strong systems design expertise.
- 8+ years of experience in Data Engineering and Platform Engineering.
- Expert-level proficiency in Java and backend engineering.
- Strong experience developing high-scale distributed systems and microservices.
2. Data Engineering :
- Extensive experience with :
1. Apache Spark
2. Apache Flink
3. Hadoop Ecosystem
4. Airflow
5. Databricks
6. Delta Lake / Iceberg / Hudi
- Expertise in ETL/ELT architecture and metadata-driven data pipelines.
3. Kafka & Streaming Technologies :
- Deep expertise in :
1. Apache Kafka
2. Kafka Streams
3. Kafka Connect
4. Schema Registry
5. Event-Driven Architecture
- Experience designing enterprise-grade streaming platforms.
4. Distributed Systems :
- Strong understanding of :
1. CAP Theorem
2. Consistency Models
3. Distributed Transactions
4. Consensus Algorithms
5. Load Balancing
6. High Availability Architectures
- Experience designing systems processing billions of events and large-scale datasets.
5. Artificial Intelligence & Modern Data Platforms :
- Strong understanding of :
1. Machine Learning Platforms
2. Generative AI
3. LLM Ecosystems
4. RAG Architectures
5. Vector Databases
6. AI Data Pipelines
- Experience enabling AI workloads through enterprise data platforms.
6. Cloud & Platform Engineering :
- Expertise in one or more cloud platforms :
1. AWS
2. Azure
3. Google Cloud
- Experience with :
1. Kubernetes
2. Docker
3. Infrastructure as Code
4. CI/CD Pipelines
5. Platform Observability
The job is for:
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1663398