HamburgerMenu
hirist

Senior DevOps Engineer - Kubernetes

WorkCrew
6 - 12 Years
Multiple Locations

Posted on: 21/09/2026

Job Description

Role Overview :


We are looking for a hands-on Senior DevOps Engineer to own end-to-end customer deployments across cloud, Bring Your Own Cloud (BYOC), air-gapped, and data-center environments. This is a high-impact, customer-facing role where you will troubleshoot complex infrastructure, maintain robust observability, and drive customer escalations to resolution with a "fix first, optimize later" mindset.


Key Responsibilities :


- Own and execute end-to-end customer deployments across cloud, BYOC, air-gapped, and data-center environments.


- Deploy, operate, and scale production environments using Kubernetes, Helm, and Terraform.


- Troubleshoot complex Kubernetes, networking, infrastructure, application, and deployment issues.


- Build and maintain comprehensive observability, monitoring, alerting, and reliability frameworks.


- Handle customer escalations with urgency and drive issues to quick resolution.


- Automate deployment and operational workflows using Python or Golang.


- Collaborate closely with Engineering, Product, and Customer Success teams to resolve production challenges.


- Participate in on-call and shift rotations, including critical customer escalations outside standard working hours.


Candidate Requirements :


Mandatory Requirements :


- Experience: 6+ years of hands-on DevOps experience deploying, operating, troubleshooting, and scaling enterprise B2B SaaS environments.


- Deployment Expertise: Proven experience owning end-to-end customer deployments across cloud and enterprise/restricted environments.


- Kubernetes: Strong hands-on experience covering troubleshooting, workloads, networking, storage, RBAC, and security.


- Helm: Practical experience with Helm chart deployment, templating, and troubleshooting.


- Terraform: Hands-on experience with infrastructure provisioning and automation.


- Linux & Cloud: Strong Linux operating system skills and cloud infrastructure troubleshooting abilities.


- Observability: Practical experience with metrics, logs, traces, and alerting systems.


- Cloud Platforms: Hands-on experience with at least one major cloud provider (AWS, GCP, Azure, or OCI).


- Mindset: Strong incident-management and production-troubleshooting mindset with a "fix first, optimize later" approach.


Preferred Requirements :


- Prior experience with BYOC, air-gapped, or restricted-network enterprise deployments.


- Familiarity with observability tools like Prometheus, Grafana, OpenTelemetry, ELK, OpenSearch, or Datadog.


- Experience with GitOps, CI/CD pipelines, and scripting in Python or Golang.


- Exposure to multi-tenant SaaS scaling, SRE concepts (SLIs/SLOs, DR, High Availability), or AI/ML pipeline deployments / iPaaS connectors.


info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...