HamburgerMenu
hirist

Job Description

The role :

Tripstack is moving its entire data stack from bare-metal VMs to Kubernetes on OpenStack in a new data centre. We are looking for a senior infrastructure engineer who has done stateful migrations before, who can plan, codify, and execute this one safely, and who will own the day-to-day operational health of the data platform once we are there.


This is a hands-on platform and SRE role with a clear, time-bounded mission. You will partner closely with our SRE team on networking, hardware, and Kubernetes fundamentals. You will own the data end to end on the data applications Druid, Spark, Redpanda, Airflow, PostgreSQL, and Elasticsearch.


This is not a machine learning role. We have a separate plan for evolving our MLFlow platform, and the right hire here may grow into more of that work over time, but day-one impact is the migration and the operational health of the platform.


Responsibilities :

Lead the data-stack migration :

- Plan and execute the migration of Druid, Spark, Redpanda, and our orchestration layer from bare-metal VMs to Kubernetes on OpenStack, with no downtime on stateful workloads.

- Design StatefulSet, PVC, pod-disruption-budget, and rolling-upgrade patterns that are safe for production data systems.

- Codify the migration with Infrastructure as Code - Terraform for OpenStack, Helm or Kustomize for Kubernetes, GitOps via ArgoCD or Flux - so the result is reproducible and supportable by the whole team.

Operate the data platform :

- Own the operational health of our production data systems, including Druid, Spark, Redpanda, Airflow, PostgreSQL, and Elasticsearch.


- You will handle segment lifecycle, JVM tuning, ingestion specs, broker/coordinator/overlord internals, partition design, consumer lag, and replication tuning.

- Build the KPIs, alerting, dashboards, and runbooks that let us see cluster exhaustion before it becomes an incident, and diagnose it quickly when it does.

- Own the query, report, segment, and tiering optimisations that keep our analytics cost-effective and responsive under load.

Raise the bar on observability and reliability to end :

- Build the Prometheus, Grafana, and distributed-tracing coverage our data systems need. Treat SLOs, error budgets, and post-incident discipline as table stakes.

- Partner with SRE on hardware, networking, and Kubernetes fundamentals, while owning the data applications themselves end to end.

Requirements :

- Strong Kubernetes experience with stateful workloads - StatefulSets, PVCs, pod disruption budgets, and rolling upgrades for data systems. You have done a real stateful migration before and can talk through what went wrong.

- Infrastructure as Code at a senior level - Terraform, Helm or Kustomize, GitOps with ArgoCD or Flux. You have shipped production infrastructure this way, not just experimented with it.

- Observability and RCA discipline - Prometheus, Grafana, distributed tracing, SLOs, error budgets, and the habit of writing the runbook that stops the next incident.

- Production operations experience with at least one of Apache Druid, Apache Kafka or Redpanda, Apache Spark, or Elasticsearch - deep enough to be credible on internals and willing to learn the others.

- 7+ years building and operating production data or platform systems, at least 2 of them on self-hosted or bare-metal infrastructure. You have been on-call for what you built.

- Clear written and verbal English; comfortable working across Krak- w, Toronto, Pune, and Stockholm time zones.

Strong differentiators :

- Security and secrets management in Kubernetes - Vault, network policies, encryption at rest.

- Change Data Capture patterns, especially PostgreSQL-to-Druid streaming.

- Delta Lake or Apache Iceberg experience and architectural judgement about when to introduce a table format.

- Exposure to travel, flights, or large-scale search and cache systems.

- Interest in growing into ownership of our MLFlow-based ML platform over time. We have a separate plan for that work, but a curious operator is welcome.

Additional Experience That Would Be Considered An Asset :

- OpenStack familiarity: Neutron networking, Cinder/Ceph storage.

- Delta Lake or Apache Iceberg experience and architectural judgement about when to introduce a table format.

- Change Data Capture patterns, especially PostgreSQL-to-Druid streaming.

- Security and secrets management in K8s - Vault, network policies, encryption at rest.

- Exposure to travel, flights, or large-scale search and cache systems.

- Experience with experimentation platforms (e.g. Growthbook, internal A/B frameworks).

Interview Process :

- Future Skill Assessment (Currency Exchange Test - Coding Round) - 60 min.

- Techno-Managerial Round - 60 min.

- Techno-Managerial Round with Architect - 60 min.

- Techno-Managerial Round with CTO - 60 min.

Working Model :

- Location: Hybrid (2 days in-office per week).

- Schedule: Monday to Friday, 12:00 PM to 9:00 PM IST (ensuring necessary overlap with global stakeholders).

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...