Posted on: 09/10/2026
Role Overview :
Owns the pipelines, data models, and analytical output built on top of the platform.
Core Responsibilities :
- Designs and builds real-time CDC pipelines, Flink stream processing jobs, and ClickHouse data models that deliver factory analytics, dashboard data, and agentic integration endpoints.
- Consumes the platform provided by the Platform Engineer.
Tech Stack :
- Streaming - CDC : Apache Kafka (producers, consumers, offsets), Kafka Connect (source/sink, SMTs), Debezium CDC (relational database source connectors, event schemas, slot management).
- Stream Processing : Apache Flink (DataStream API, Table API/SQL, stateful processing, windowing, joins), PyFlink or Flink SQL, real-time aggregation, enrichment, deduplication, late-data handling, and Complex Event Processing (CEP).
- Databases - Data Modeling : ClickHouse (table engine selection, materialized views, TTL policies, query optimization), SQL proficiency, time-series and analytical data modeling (wide tables, pre-aggregation, partitioning), Schema Registry (Avro, JSON Schema).
- Python Development : Pipeline development, data transformation, pandas, PyArrow, SQLAlchemy, Kafka/Flink Python clients, FastAPI or Flask for data service APIs, unit/integration testing.
- AI-Assisted Development : Proficiency with AI coding assistants (GitHub Copilot, Cursor, Claude), scaffolding Flink jobs, generating ClickHouse DDL, prompt engineering, and critical evaluation of AI-generated code.
- Domain - Integration Skills : Factory data source integration (MES, SCADA, factory tool messaging), SECS/GEM, OPC-UA, or MQTT (preferred), dashboard integration (Apache Superset, Grafana, Tableau, or Power BI), and agentic/LLM integration patterns.
Did you find something suspicious?
Posted by
Posted in
Data Engineering
Functional Area
Data Engineering
Job Code
1677574