Posted on: 17/06/2026
Job Title : Data Platform Architect Engineer (Hands-On Builder)
Location : Remote
Role Overview :
This is a foundational, high-impact role responsible for designing and building the core data and AI platform capabilities for our startup and client environments. We are looking for a deeply hands-on Data Architect with elite engineering skills, someone who can define architecture, write production-grade code, and create reusable accelerators that scale across clients. This is not a pure architect role.
This is not a pure engineering role. This is a builder-architect role, you will own the platform from zero to scale. You will design and implement modern data platforms (warehouse, lakehouse, streaming, and semantic layers) and directly build the components required to operationalize them.
Key Responsibilities :
1. End-to-End Data Platform Architecture & Build :
- Design and implement data warehouses, lakehouses, and hybrid architectures.
- Build batch, micro-batch, and real-time streaming pipelines.
- Architect and implement event-driven and API-based ingestion frameworks.
- Define and build canonical data models and enterprise data architecture patterns.
2. Deep Engineering Execution (Hands-On Coding Required) :
- Write production-grade code in Python, SQL, and/or Scala.
- Build data pipelines, transformation frameworks, and reusable components.
- Develop platform accelerators, templates, and SDK-like assets.
- Implement CI/CD for data pipelines and infrastructure-as-code.
3. Data Integration & Activation (Critical Emphasis) :
- Design and build data ingestion frameworks (ETL/ELT, CDC, APIs, streaming).
- Implement data activation layers for downstream consumption (APIs, reverse ETL, applications).
- Enable real-time and batch data serving patterns.
- Build integrations across enterprise systems (ERP, CRM, SaaS, operational systems).
4. Semantic, Metadata & Context Layer :
- Design and implement :
1. Semantic / metrics layer.
2. Metadata management and data catalog.
3. Data lineage and observability frameworks.
4. Data context layer (business-friendly abstraction of data).
- Enable self-service analytics and AI-ready data foundations.
5. Lakehouse & Modern Data Stack Implementation :
- Build and operationalize :
1. Lakehouse architectures (Delta Lake / Iceberg / Hudi).
2. Data warehouse platforms.
3. Streaming platforms.
- Optimize for performance, scalability, cost, and reliability.
6. Platform Engineering & Infrastructure :
- Configure and optimize cloud-native data platforms.
- Implement infrastructure-as-code (Terraform, etc.).
- Design multi-tenant, reusable platform architectures.
- Ensure security, governance, and compliance (PII, PHI, etc.).
7. Reusable Assets & Startup Acceleration :
- Build reusable accelerators, frameworks, and reference architectures.
- Create plug-and-play modules for rapid client deployment.
- Contribute to IP creation (toolkits, patterns, templates).
Required Skills & Experience :
- Strong hands-on experience in data engineering and data architecture.
- Expertise in Python, SQL, and/or Scala.
- Experience with modern data platforms (lakehouse, warehouse, streaming).
- Deep understanding of ETL/ELT, CDC, APIs, and event-driven systems.
- Experience with cloud platforms (AWS, Azure, or GCP).
- Familiarity with Terraform or other infrastructure-as-code tools.
- Strong understanding of data modeling, governance, and security.
- Experience building scalable, production-grade data systems.
What We Are Looking For :
- Builder mindset with strong ownership.
- Ability to operate from zero to scale.
- Balance of architecture thinking and hands-on execution.
- Passion for creating reusable, scalable solutions.
- Startup mindset with high adaptability and speed.
Why Join Us :
- Opportunity to build core platform capabilities from scratch.
- High ownership and impact role.
- Work on cutting-edge data and AI systems.
- Shape foundational architecture across clients and products.
Skills & Qualifications :
Must-Have Skills :
Tier 1 (Non-Negotiable) :
- Expert-level SQL (query optimization, large-scale data processing).
- Strong programming skills : Python (mandatory), along with Scala or Java.
- Deep hands-on experience with :
1. Snowflake.
2. Databricks.
3. Amazon Web Services (AWS) or equivalent (Azure / GCP).
- Strong experience in data warehousing and lakehouse implementations.
- Proven experience in building end-to-end data pipelines (batch and streaming).
- Expertise in data integration (APIs, CDC, event streams, file-based ingestion).
- Hands-on experience in real-time data processing (Kafka / streaming ecosystems).
- Experience with CI/CD and DevOps practices for data platforms.
Tier 2 (Advanced / Architecture Depth) :
- Experience with lakehouse frameworks: Delta Lake, Apache Iceberg, Apache Hudi.
- Orchestration tools : Airflow, Dagster, or equivalent.
- Stream processing frameworks : Spark Streaming, Flink, Kafka Streams.
- Knowledge of data contracts, schema registry, and governance frameworks.
- Experience with metadata management, data catalog, and lineage systems.
- Implementation of semantic / metrics layers.
- Infrastructure-as-code tools: Terraform, CloudFormation.
- Experience designing multi-cloud and hybrid data architectures.
Nice-to-Have Skills :
- Experience with data mesh and domain-oriented architecture.
- Familiarity with data observability tools (Monte Carlo, Datadog, etc.).
- Exposure to enterprise platforms such as :
1. SAM (Software Asset Management).
2. TAM (Technology Asset Management).
3. HAM (Hardware Asset Management).
4. ITSM (IT Service Management).
5. CMDB (Configuration Management Database).
- Exposure to AI/ML data pipelines and feature stores.
- Knowledge of reverse ETL and operational analytics.
Soft Skills & Mindset :
- Builder-first mindset (focus on execution, not just design).
- Strong ownership and accountability.
- Ability to operate in 0-1 startup environments.
- Excellent problem-solving and system thinking capabilities.
- Comfortable working in ambiguous environments and directly with clients.
Industry Exposure :
- Experience across two or more of the following industries is preferred : Technology, Financial Services, Healthcare, Retail, Manufacturing, Energy, Media, Government, Education, Hospitality.
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1645590