Posted on: 21/09/2026
Roles & Responsibilities :
- Own end-to-end customer deployments across cloud, BYOC, air-gapped, and data-center environments.
- Deploy and operate production environments using Kubernetes, Helm, and Terraform.
- Troubleshoot complex Kubernetes, networking, infrastructure, application, and deployment issues.
- Build and maintain observability, monitoring, alerting, and reliability.
- Handle customer escalations and drive issues to resolution with a "fix first, optimize later" mindset.
- Automate deployment and operational workflows.
- Work closely with Engineering, Product, and Customer teams to resolve production challenges.
- Participate in on-call and shift rotations, including critical customer escalations outside standard working hours.
Tech Stack & Requirements :
- 6+ years of hands-on DevOps experience, deploying, operating, troubleshooting, and scaling enterprise SaaS environments.
- Strong hands-on Kubernetes experience : troubleshooting, networking, workloads, storage, RBAC, and security.
- Hands-on Helm experience : deployment, templating, and troubleshooting.
- Hands-on Terraform experience : infrastructure provisioning and automation.
- Strong Linux and cloud infrastructure troubleshooting skills.
- Hands-on experience with observability : metrics, logs, traces, and alerting.
- Hands-on experience with one or more of AWS / GCP / Azure / OCI.
- Strong incident-management and production-troubleshooting mindset.
- Experience with B2B SaaS product companies.
- Preferred : BYOC, air-gapped, or restricted-network environments; Prometheus, Grafana, OpenTelemetry, ELK, OpenSearch, Datadog; GitOps, CI-CD; Python or Golang; AI/ML pipeline deployments; SRE concepts.
The job is for:
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
DevOps / Cloud
Job Code
1673026