Posted on: 25/05/2026
Description :
Role Overview :
- Hands-on position focused on Kubernetes, AWS (mandatory), GCP, and Python automation
- Design and maintain highly available Kubernetes systems; define/enforce SLOs, SLIs, error budgets
- Lead incident response, RCA, and postmortems; drive automation for reliability improvements
- Architect observability platforms using Prometheus, Grafana, OpenTelemetry, GCP Monitoring; set actionable alerting standards
- Operate GKE/EKS clusters; manage containerized workloads with Docker; deploy services via Helm
- Develop Python-based automation and reliability tools; integrate CI/CD pipelines with observability checks
- Mentor junior engineers; influence architecture decisions; collaborate across engineering teams
Candidate Profile :
- Bachelors degree in Computer Science, IT, Engineering, or related field (Masters preferred)
- 6-10 years in Site Reliability Engineering or related roles
- Strong Kubernetes (GKE/EKS) expertise
- Mandatory AWS experience; GCP exposure preferred
- Skilled in Python automation and CI/CD integration
- Experience with observability tools (Prometheus, Grafana, OpenTelemetry)
- Proven track record in incident response, RCA, and postmortems
- Ability to mentor engineers and influence architecture decisions
Did you find something suspicious?
Posted by
Recruiter
Last Active: NA as recruiter has posted this job through third party tool.
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1638677