Posted on: 05/08/2026
Job Description :
We are looking for an Apache Doris Developer who can own the MPP query engine, lead migration of existing pipelines and warehouses, and help establish a scalable, low-latency real-time analytics (OLAP) platform.
You will work at the intersection of query engineering and data modeling - translating legacy Snowflake/Spark logic into performant Doris workloads (both native tables and Iceberg external catalogs) while ensuring correctness, parity, and reliability.
Key Responsibilities :
1. Migration & Delivery :
- Migrate SQL workloads, transformations, and data models from Snowflake and Spark to Apache Doris.
- Translate Snowflake-specific SQL (semi-structured VARIANT / OBJECT, stored procedures, tasks, streams) and Spark jobs into equivalent, validated Doris SQL.
- Design Doris table models appropriately - Duplicate, Aggregate, and Unique key models - mapping source schemas to the right model for each workload.
- Integrate Apache Iceberg via Doris Multi-Catalog for federated lakehouse querying, and design ingestion paths from Iceberg into native Doris tables where low latency is required.
- Design and implement data-parity validation - schema, row-count, and value-level reconciliation between source (Snowflake/Spark) and target (Doris/Iceberg).
2. Query Engineering & Performance :
- Write, optimize, and troubleshoot complex analytical SQL on Doris (MySQL-protocol compatible).
- Tune query performance : bucketing/partitioning strategy, colocate joins, runtime filters, and the cost-based optimizer.
- Design and maintain materialized views, rollups, and indexes (inverted, bitmap, bloom filter, N-gram) to accelerate queries.
- Analyze query plans (EXPLAIN, profile) and resolve memory pressure, tablet skew, and slow scans.
Required Skills & Qualifications :
- Hands-on production experience with Apache Doris (or a comparable MPP OLAP engine - StarRocks, ClickHouse, Greenplum).
- Strong, deep SQL expertise - complex analytical queries, window functions, CTEs, query optimization.
- Solid understanding of Doris architecture (FE/BE), the three data models (Duplicate/Aggregate/Unique), partitioning, bucketing, and tablet/replica management.
- Experience with Apache Iceberg (or comparable open table formats - Delta Lake, Hudi) and lakehouse / external-catalog federation.
- Practical experience with Snowflake and/or Apache Spark - enough to read, understand, and migrate existing workloads.
- Understanding of distributed / MPP query execution : join distribution, runtime filters, memory management, and data skew.
- Experience with cloud object storage and columnar file formats (Parquet, ORC).
- Proficiency in at least one programming language (Python, Java, or Scala) for tooling, UDFs, and automation.
- Version control (Git) and CI/CD for data pipelines.
Preferred / Nice-to-Have :
- Experience leading a Snowflake - Doris or Spark - Doris migration at scale.
- Experience building data-validation / reconciliation frameworks for migration parity.
- Knowledge of Doris internals or connector development (custom UDFs/UDAFs).
- Workflow orchestration (Airflow, Dagster) and dbt (with a Doris adapter).
- Data governance, lineage, and cost-optimization tooling.
Soft Skills :
- Strong analytical and problem-solving mindset for debugging correctness and performance issues.
- Clear communication - able to document migration decisions and work with analytics, platform, and business teams.
- Ownership mentality with attention to data correctness and reliability.
Did you find something suspicious?
Posted by
Ruchi
VP- Recruitment and Operations at CONVERGENCE CONSULTING PRIVATE LIMITED
Last Active: 11 Aug 2026
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1660766