Posted on: 14/05/2026
Description :
- Manage and operate Kubernetes clusters including deployments, scaling, and upgrades
- Support tenant onboarding, provisioning, and environment configuration
- Assist in infrastructure provisioning using Terraform
- Deploy and manage services using Helm charts
- Monitor system performance using logs, metrics, and traces (MELT)
- Troubleshoot production issues across applications, infrastructure, networking, and databases
- Support CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI)
- Maintain observability dashboards, alerts, and operational runbooks
- Assist in backup strategies, disaster recovery, and high availability setups
- Participate in incident response and root cause analysis (RCA), documenting learnings
Required Skills & Qualifications :
- Basic understanding of Linux systems, networking concepts, and cloud platforms (AWS/Azure/GCP)
- Working knowledge of Kubernetes and Docker
- Hands-on exposure to :
a. Terraform (modules, state management, plan/apply)
b. Helm (charts, values.yaml, releases)
- Familiarity with monitoring tools such as Prometheus, Grafana, or Datadog
- Basic scripting skills (Python or Bash)
- Understanding of CI/CD pipelines and DevOps practices
Good to Have :
- Exposure to distributed systems and microservices architecture
- Understanding of multi-tenant architectures and isolation strategies
- Familiarity with databases (PostgreSQL, Redis) and messaging systems (Kafka/Pulsar)
- Basic knowledge of security practices (IAM, secrets management, TLS)
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1635785