HamburgerMenu
hirist

Platform Engineer - Kafka/ClickHouse

Dhruv Compusoft Consultancy
3 - 6 Years
Bangalore

Posted on: 09/10/2026

Job Description

Core Responsibility Summary:

Deploys, operates, and evolves the Kafka - Flink - ClickHouse platform on Kubernetes across multiple factory data centers.

Ensures the platform is reliable, observable, performant, and ready for data engineering teams to build on top of.

Kubernetes & Container Platform:

- Deep Kubernetes expertise (workload types, scheduling, resource requests/limits, namespaces, RBAC).

- Kubernetes operators for stateful systems - Strimzi (Kafka), Flink Kubernetes Operator, ClickHouse Operator (Altinity).

- Helm chart authoring and management (not just consumption - writing and maintaining charts).

- StatefulSet management, persistent volume claims, storage class configuration.

- Multi-cluster and multi-site Kubernetes deployments.

- Node affinity, taints/tolerations, pod disruption budgets for workload isolation.

- Kubernetes networking: services, ingress, network policies, DNS.

- Secrets management (Kubernetes secrets, Vault integration, or equivalent).

Platform Component Expertise :

Kafka :

- Kafka cluster sizing, broker configuration, topic design (partitions, replication, retention).

- Strimzi operator - KafkaNodePool, KafkaTopic, KafkaUser, KafkaMirrorMaker2 CRDs.

- Kafka Connect cluster deployment and connector lifecycle management.

- Kafka security: TLS, mTLS, SASL, ACLs.

- Kafka upgrade and rolling restart procedures with zero data loss.

Flink :

- Flink cluster deployment modes on Kubernetes (Application mode, Session mode).

- Flink Kubernetes Operator - FlinkDeployment, FlinkSessionJob CRDs.

- Checkpointing and savepoint management (S3 or distributed storage backends).

- Flink HA configuration (Kubernetes-native HA or ZooKeeper).

- Job submission, monitoring, and restarts without data loss.

ClickHouse :

- ClickHouse cluster deployment with Altinity operator or ClickHouse Operator.

- Shard and replica topology design for multi-factory deployments.

- Storage configuration: tiered storage, disk policies, S3-backed cold storage.

- Backup and restore strategies.

- ClickHouse Keeper vs ZooKeeper for coordination.

Observability & Monitoring :

- Full observability stack deployment: Prometheus, Grafana, Alertmanager.

- Kafka metrics: broker JMX metrics, consumer lag (Burrow or Cruise Control), under-replicated partitions.

- Flink metrics: checkpoint duration, backpressure, restart frequency, throughput/latency per operator.

- ClickHouse metrics: query latency, merge backlog, insert rate, replication lag, memory/disk pressure.

- Kubernetes cluster metrics: node utilization, pod restarts, OOMKill events, PVC utilization.

- SLO/SLI definition and alerting for platform uptime and data freshness.

- Distributed tracing for pipeline end-to-end latency visibility.

- Log aggregation (EFK stack, Loki, or equivalent) for platform component logs.

Load Testing & Capacity Planning:

- Load testing frameworks for streaming platforms (Gatling, k6, custom Kafka producer harnesses).

- Designing load test scenarios that simulate factory production volumes and burst conditions.

- Benchmarking Kafka throughput (MB/s, msg/s) under varying partition counts and consumer configurations.

- Flink job stress testing - measuring backpressure onset, checkpoint interval degradation under load.

- ClickHouse ingestion benchmarking (insert throughput, concurrent query performance under load).

- Capacity modeling: translating factory data volumes and growth projections into infrastructure sizing.

- Breaking point analysis - identifying bottlenecks before production traffic hits them.

Isolation & Multi-Tenancy:

- Namespace-level isolation for multiple factory deployments on shared Kubernetes clusters.

- Kafka topic and consumer group isolation strategies across factory sites and teams.

- Resource quotas and LimitRanges to prevent noisy-neighbor problems between factory workloads.

- Network policy enforcement to isolate factory data plane traffic.

- Separate Kafka clusters vs shared clusters with namespace isolation - trade-off analysis and implementation.

- Per-factory ClickHouse cluster deployment vs shared cluster with database/user isolation.

CI/CD & GitOps for Platform:

- GitOps-driven platform deployment: ArgoCD or Flux for Kubernetes manifest management.

- Helm chart versioning and promotion across environments (dev - staging - factory-prod).

- Operator upgrade management and CRD migration procedures.

- Automated smoke tests post-deployment to validate platform health.

- Change management for platform updates in production factory environments.

Security & Compliance:

- mTLS and certificate management (cert-manager on Kubernetes).

- Kafka ACL management at scale.

- Network segmentation between factory OT and IT data flows.

- Image scanning and supply chain security for platform container images.

- Audit logging for platform access and configuration changes.

AI-Assisted Development:

- Using AI coding assistants (GitHub Copilot, Cursor, or equivalent) to accelerate Helm chart authoring, operator configuration, and runbook generation.

- AI-assisted troubleshooting - using LLMs to diagnose Kubernetes events, Flink exceptions, and Kafka error logs faster.

- Generating observability dashboards and alert rules with AI assistance.

- Critical review of AI-generated infrastructure code before applying to production.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...