Posted on: 09/10/2026
Core Responsibility Summary:
Deploys, operates, and evolves the Kafka - Flink - ClickHouse platform on Kubernetes across multiple factory data centers.
Ensures the platform is reliable, observable, performant, and ready for data engineering teams to build on top of.
Kubernetes & Container Platform:
- Deep Kubernetes expertise (workload types, scheduling, resource requests/limits, namespaces, RBAC).
- Kubernetes operators for stateful systems - Strimzi (Kafka), Flink Kubernetes Operator, ClickHouse Operator (Altinity).
- Helm chart authoring and management (not just consumption - writing and maintaining charts).
- StatefulSet management, persistent volume claims, storage class configuration.
- Multi-cluster and multi-site Kubernetes deployments.
- Node affinity, taints/tolerations, pod disruption budgets for workload isolation.
- Kubernetes networking: services, ingress, network policies, DNS.
- Secrets management (Kubernetes secrets, Vault integration, or equivalent).
Platform Component Expertise :
Kafka :
- Kafka cluster sizing, broker configuration, topic design (partitions, replication, retention).
- Strimzi operator - KafkaNodePool, KafkaTopic, KafkaUser, KafkaMirrorMaker2 CRDs.
- Kafka Connect cluster deployment and connector lifecycle management.
- Kafka security: TLS, mTLS, SASL, ACLs.
- Kafka upgrade and rolling restart procedures with zero data loss.
Flink :
- Flink cluster deployment modes on Kubernetes (Application mode, Session mode).
- Flink Kubernetes Operator - FlinkDeployment, FlinkSessionJob CRDs.
- Checkpointing and savepoint management (S3 or distributed storage backends).
- Flink HA configuration (Kubernetes-native HA or ZooKeeper).
- Job submission, monitoring, and restarts without data loss.
ClickHouse :
- ClickHouse cluster deployment with Altinity operator or ClickHouse Operator.
- Shard and replica topology design for multi-factory deployments.
- Storage configuration: tiered storage, disk policies, S3-backed cold storage.
- Backup and restore strategies.
- ClickHouse Keeper vs ZooKeeper for coordination.
Observability & Monitoring :
- Full observability stack deployment: Prometheus, Grafana, Alertmanager.
- Kafka metrics: broker JMX metrics, consumer lag (Burrow or Cruise Control), under-replicated partitions.
- Flink metrics: checkpoint duration, backpressure, restart frequency, throughput/latency per operator.
- ClickHouse metrics: query latency, merge backlog, insert rate, replication lag, memory/disk pressure.
- Kubernetes cluster metrics: node utilization, pod restarts, OOMKill events, PVC utilization.
- SLO/SLI definition and alerting for platform uptime and data freshness.
- Distributed tracing for pipeline end-to-end latency visibility.
- Log aggregation (EFK stack, Loki, or equivalent) for platform component logs.
Load Testing & Capacity Planning:
- Load testing frameworks for streaming platforms (Gatling, k6, custom Kafka producer harnesses).
- Designing load test scenarios that simulate factory production volumes and burst conditions.
- Benchmarking Kafka throughput (MB/s, msg/s) under varying partition counts and consumer configurations.
- Flink job stress testing - measuring backpressure onset, checkpoint interval degradation under load.
- ClickHouse ingestion benchmarking (insert throughput, concurrent query performance under load).
- Capacity modeling: translating factory data volumes and growth projections into infrastructure sizing.
- Breaking point analysis - identifying bottlenecks before production traffic hits them.
Isolation & Multi-Tenancy:
- Namespace-level isolation for multiple factory deployments on shared Kubernetes clusters.
- Kafka topic and consumer group isolation strategies across factory sites and teams.
- Resource quotas and LimitRanges to prevent noisy-neighbor problems between factory workloads.
- Network policy enforcement to isolate factory data plane traffic.
- Separate Kafka clusters vs shared clusters with namespace isolation - trade-off analysis and implementation.
- Per-factory ClickHouse cluster deployment vs shared cluster with database/user isolation.
CI/CD & GitOps for Platform:
- GitOps-driven platform deployment: ArgoCD or Flux for Kubernetes manifest management.
- Helm chart versioning and promotion across environments (dev - staging - factory-prod).
- Operator upgrade management and CRD migration procedures.
- Automated smoke tests post-deployment to validate platform health.
- Change management for platform updates in production factory environments.
Security & Compliance:
- mTLS and certificate management (cert-manager on Kubernetes).
- Kafka ACL management at scale.
- Network segmentation between factory OT and IT data flows.
- Image scanning and supply chain security for platform container images.
- Audit logging for platform access and configuration changes.
AI-Assisted Development:
- Using AI coding assistants (GitHub Copilot, Cursor, or equivalent) to accelerate Helm chart authoring, operator configuration, and runbook generation.
- AI-assisted troubleshooting - using LLMs to diagnose Kubernetes events, Flink exceptions, and Kafka error logs faster.
- Generating observability dashboards and alert rules with AI assistance.
- Critical review of AI-generated infrastructure code before applying to production.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1677582