Posted on: 27/08/2026
Position : Data Engineer
Mandatory skills : Hudi-heavy Data Engineer (or similar skill)
Years of exp : 6-10 years
Location : Bangalore
Role : Data Engineer
We are looking for a skilled Data Engineer to help us build and scale the data backbone for our next generation Connected Two-Wheeler (2W) ecosystem. In this role, you will design, ingest, process, and store high-velocity IoT telemetry data from hundreds of thousands of vehicles on the road. Working at the intersection of streaming data, batch processing, transactional data lakes, and modern enterprise data warehousing, your work will directly enable predictive maintenance, rider safety features, and deep R&D insights.
Responsibilities :
1. Data Ingestion, Streaming & Search Analytics :
- Batch & Streaming Pipelines : Build and manage data ingestion pipelines using AWS Glue, extracting data from multiple sources - including Kafka, OpenSearch, and PostgreSQL - and routing it reliably into our Data Lake and Data Warehouse.
- Search & Near-Real-Time Analytics : Implement and maintain OpenSearch indices to enable rapid log analysis, anomaly detection, and vehicle-state troubleshooting for engineering and support teams.
2. Transactional Data Lake & Lakehouse Engineering :
- Medallion Architecture : Architect and maintain a structured Bronze / Silver / Gold (Medallion) data model to cleanly separate raw IoT ingestion, validated/curated datasets, and business-ready analytics tables.
- Apache Hudi Architecture : Build, manage, and scale our ACID-compliant Data Lake on AWS S3 using Apache Hudi.
- Metadata & Optimization : Deeply leverage Hudi metadata tables, the timeline server, and smart indexing strategies (Bloom/Record-level) to ensure high-performance upserts, Change Data Capture (CDC), and clean schema evolution as vehicle payloads change.
3. Distributed Compute & Orchestration :
- PySpark & Glue Engineering : Write clean, highly scalable ETL/ELT applications using PySpark, deployed across AWS EMR and AWS Glue.
- Performance Tuning : Optimize Spark logic for memory management, partitioning, skew handling, and serialization when crunching multi-terabyte vehicle telemetry logs.
- Workflow Orchestration : Design robust, end-to-end scheduled and event-driven data workflows using AWS Step Functions and AWS Glue Workflows.
4. Data Warehousing & Operational Databases :
- Enterprise Data Warehousing : Own the architecture, data modeling, and performance tuning in Snowflake to support downstream BI, connected-fleet dashboards, and data science workloads.
- Relational Systems : Manage extraction, replication, and state-tracking across operational PostgreSQL (and Amazon RDS/Aurora) databases.
5. Automotive Domain & Data Governance :
- Telemetry & Time-Series Modeling : Build specialized schemas tailored for high-frequency time-series data to accurately model and compare vehicle behavior across both EV and ICE architectures.
- Data Governance & Observability : Implement automated data quality frameworks (e.g., Great Expectations, Deequ) to catch stale telemetry, corrupted packets, or GPS drift early.
- FinOps & Cost Optimization : Put smart data retention and tiering strategies in place (S3 Hot/Cold storage, Hudi compaction, Snowflake clustering) to keep cloud storage and compute costs lean as our fleet scales.
What you'll need :
Relevant Experience :
- 6 - 10 years
Required Technical Stack & Qualifications :
- Distributed Compute : Apache Spark, PySpark, Spark SQL, EMR, AWS Glue
- Data Lake & Metadata : Apache Hudi (Strong hands-on understanding of Hudi metadata tables, compaction, clustering, schema evolution, and ACID properties on S3)
- Streaming & Search : Apache Kafka / AWS MSK, OpenSearch / Elasticsearch
- Data Warehousing & RDBMS : Snowflake, PostgreSQL, Amazon S3
- Cloud & Orchestration : AWS Ecosystem (Step Functions, IAM, CloudWatch, EMR, Glue, S3)
- Languages : Python, SQL
- Domain Knowledge : Experience working with IoT / Telemetry / Time-series data (Automotive 2W/4W or EV Battery Management Systems is a massive plus)
Behavioral Skills :
- Strategic Thinker : Ability to translate details into bigger picture implications driving the business forward, challenging the status quo. Aligns the right resources to the task at hand; foresees and plans around obstacles.
- Talent Management : Has a passion for building great teams - proven ability to develop, motivate and champion talent beyond own organization.
- Innovate for Growth : Technology Evangelist. Always thinking about how to make improvements; able to implement changes that map to business strategy. Stays abreast of cutting edge technology trends.
- Lead & Adapt to Change : Thrives in a changing, dynamic environment and can drive operational efficiencies that map to changing needs.
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1666481