Posted on: 20/08/2026
Role Overview :
A senior, customer-facing SRE who owns the reliability of a client's production systems end-to-end. Embedded as the main technical point of contact, you will design reliable infrastructure, lead incidents, apply AI-assisted operations, and mentor the team.
Key Responsibilities :
- Own the reliability of production systems for one or more enterprise customers on AWS.
- Define and manage SLIs, SLOs, and error budgets, and ensure monitoring provides meaningful insight.
- Lead the response to serious (P0/P1) incidents and run blameless post-mortems that lead to fixes.
- Lead within the on-call rotation.
- Operate and optimise Kubernetes clusters and AWS services as workloads grow.
- Introduce AI SRE tooling and AIOps where it speeds up triage and resolution, with appropriate guardrails.
- Advise customers on improving their reliability practices, and mentor associate engineers.
- Share field learnings with product and engineering teams.
Required Qualifications :
- Substantial SRE experience with real ownership of production reliability on AWS.
- Experience running Kubernetes and AWS services at scale.
- A track record of defining and managing SLIs, SLOs, and error budgets.
- Strong observability practice.
- Proven incident leadership and post-mortem facilitation.
- Strong automation skills (Python, Go, or Bash).
- Hands-on understanding of AI-assisted operations, including introducing AI SRE tooling with sensible guardrails.
- AWS certification at Associate level as a minimum.
Tech Stack :
- Cloud (AWS) : EC2, EKS, ECS, Lambda, S3, RDS, VPC, IAM, CloudWatch, ELB/ALB, Route 53.
- Containers & orchestration : Docker, Kubernetes (EKS), Helm.
- Observability & monitoring : Prometheus, Grafana, Datadog, OpenTelemetry, Loki, Tempo/Jaeger, ELK/Elastic.
- Automation & scripting : Python, Go, Bash, Git.
Soft Skills :
- Confident, customer-facing communication.
- Calm, decisive incident leadership.
- Mentors and raises standards across the team.
- Sound technical judgement.
The job is for:
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1664877