Posted on: 05/08/2026
Principal Architect :
Experience : 1215 years.
Location : Pune (onsite).
About the Role :
We are looking for a Principal Architect to lead the design and development of a Lakehouse solution end to end.
This is an opportunity to architect and build a reusable data platform framework from the ground up.
You will own the technical architecture, write code for the core framework, mentor a growing engineering team, and work closely with leadership on the overall technology direction.
The ideal candidate is a hands-on architect with a proven track record of building complex frameworks and platforms using open-source big-data technologies such as Spark, Hadoop, and modern Lakehouse stacks (e.g., Iceberg, Trino, Databricks, etc).
Key Responsibilities :
1. Architecture (40%) :
- Own the end-to-end architecture of the Lakehouse platform and its reusable framework components.
- Design a metadata-driven data engineering framework built on Spark SQL and PySpark.
- Architect a pipeline engine supporting dynamic SQL generation, incremental loads, CDC, SCD implementations, and schema evolution.
- Design Kubernetes-native deployment with multi-tenancy, resource isolation, scaling strategy, high availability, and disaster recovery.
- Define API contracts, plugin/extension architecture, and integration patterns.
- Drive and document key architecture decisions (ADRs).
2. Hands-on Engineering (40%) :
- Build core framework components : pipeline engine, dynamic SQL generation, data quality checks, error handling and retry mechanisms.
- Implement observability across the platform.
- Optimize Spark job performance (partitioning, shuffle tuning, memory management, query plans).
- Establish and implement CI/CD pipelines, automated testing (unit and integration), and release practices.
- Embed data governance and security into the platform (access control, auditability, cataloging).
3. Engineering Leadership (20%) :
- Mentor and guide junior engineers; conduct design and code reviews.
- Define coding standards, engineering best practices, and documentation norms.
- Contribute to broader technical direction and team growth.
- Continuously raise the bar on engineering quality and delivery practices.
Required Experience :
- 12 to 15 years of software engineering experience overall.
- 8+ years in data platform engineering.
- 5+ years as an Architect on Spark/Hadoop-based platforms.
- Proven success building complex, reusable frameworks or platforms (not just applications) using open-source technologies.
- Expert-level proficiency in Python, PySpark, Spark SQL, and Spark internals.
- Deep knowledge of Apache Iceberg, distributed systems, Kubernetes, Docker, REST APIs, microservices, Linux, Git, and CI/CD.
Technical Stack :
- Data Platform : Apache Spark, Apache Iceberg, Apache Ozone, Apache Ranger, Trino, Kafka, Airflow.
- Platform Engineering : Kubernetes-native deployments, Helm, Terraform, Prometheus, Grafana, containerization, multi-tenant architecture, storage management, HA/DR.
- Framework Capabilities You'll Build : JSON-driven pipelines, metadata-driven ETL, dynamic SQL generation, incremental loads & CDC, SCD handling, schema evolution, data quality, lineage, plugin architecture, unit testing, performance optimization.
What We're Looking For :
- A hands-on architect who designs systems and writes the code that proves the design.
- Strong architecture and design instincts with the judgment to keep frameworks simple and extensible.
- Complex problem-solving ability and speed in picking up new technologies.
- Experience with Agile and spec-driven development.
- Clear communication able to explain architecture decisions to engineers and leadership alike.
Nice to Have :
- Contributions to open-source data projects (Spark, Iceberg, Trino, etc.).
- Experience with Databricks or comparable managed Lakehouse platforms.
- Prior experience taking a platform or framework from concept to production.
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1660808