HamburgerMenu
hirist

Principal Site Reliability Engineer

Scaling Theory Technologies
5 - 11 Years
Bangalore

Posted on: 11/04/2026

Job Description

Key Responsibilities :

- Architect multi-tenant isolation models (tenant-per-cluster, namespace, and logical isolation).

- Design control plane + data plane separation for SaaS.

- Build tenant onboarding, provisioning, and lifecycle management.

- Implement scalable ingestion pipelines (1M+ events/day, streaming via Pulsar/Kafka).

- Design multi-tenant graph architecture (Neo4j) and storage strategies.

- Ensure RBAC, tenant-level security, and data isolation (IAM, Vault, encryption).

- Optimize cost, performance, and autoscaling across tenants.

- Enable SaaS observability, billing, metering, and usage tracking.

- Work on multi-cloud deployment (AWS, Azure, GCP) with portability.

Core Skills Required :

- Strong experience in distributed systems & SaaS architectures.

- Experience in coding HelmCharts, Terraform Scripts, Bash

- Deep expertise in Kubernetes (multi-tenant patterns, operators, scaling).

- Experience with event-driven systems (Kafka / Pulsar).

- Hands-on with databases: Neo4j, Postgres, time-series/log systems.

- Knowledge of multi-tenant security models (isolation, encryption, IAM).

- Experience building control planes / platform engineering systems.

- Familiarity with Terraform / IaC and cloud networking.

- Exposure to observability (MELT) and large-scale systems.

Good to Have :

- Experience with graph-based systems or knowledge graphs.

- Exposure to AI/LLM-integrated platforms.

- Prior work on DevOps / SRE / observability platforms.

Success Criteria :


- Deliver a production-grade multi-tenant SaaS architecture.

- Support hundreds of tenants with strict isolation and high performance.

- Enable self-serve onboarding + minimal ops overhead.

- Achieve high availability, scalability, and cost efficiency.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...