Posted on: 24/08/2026
Role Summary :
We are hiring a senior Site Reliability Engineer to own the availability, scalability, performance and operability of our production platform. The estate is AWS-first with a growing Azure footprint, and is fully provisioned as code using Terraform and CloudFormation. This is a hands-on engineering role covering all pillars of SRE infrastructure, automation, observability, SLO management, incident response, resilience and security with mentoring responsibility for mid-level engineers and participation in a shared on-call rotation.
Key Responsibilities :
- Cloud Infrastructure (AWS primary, Azure secondary): Architect, build and operate scalable AWS infrastructure (VPC and connectivity, IAM, compute, storage, managed databases, serverless, multi-account governance) and own production Amazon EKS clusters end to end provisioning, upgrades, autoscaling, networking and security. Build and support the secondary Azure footprint and drive cloud cost optimisation.
- Infrastructure as Code: Design and maintain reusable Terraform modules and CloudFormation stacks across multi-account and multi-subscription environments, with remote state, versioning, drift detection and policy-as-code guardrails. Eliminate manual console changes.
- Monitoring & Observability: Own the metrics, logs and tracing stack; build symptom- and SLO-burn-based alerting; instrument services with development teams; and maintain golden dashboards and runbooks that any on-call engineer can use.
- SLIs, SLOs & Error Budgets: Define SLIs for critical user journeys, agree SLOs with stakeholders, operate error-budget policy to balance feature velocity against reliability, and report reliability to leadership through regular service reviews.
- Automation & Toil Reduction: Measure and engineer away repetitive operational work; build tooling in Python, Bash or Go for provisioning, remediation, patching and diagnostics; and implement self-healing and auto-remediation for known failure modes.
- Incident Response & On-Call: Participate in a compensated on-call rotation, act as Incident Commander for high-severity events, run blameless postmortems with corrective actions tracked to closure, and drive down MTTD and MTTR.
- CI/CD & Release Engineering: Build and harden application and infrastructure pipelines; enable blue-green, canary and progressive delivery with automated health gates and fast rollback; and improve DORA metrics including change failure rate.
- Backup, DR & Business Continuity: Own the backup estate across Azure Backup (Recovery Services vaults, retention, immutability, cross-region restore) and AWS Backup; define RPO/RTO per service; and prove recoverability through scheduled restore tests and DR failover drills.
- Capacity Planning & Performance: Forecast demand and plan capacity across compute, storage, network and database tiers; run load and stress testing to validate headroom and autoscaling; tune performance and practise chaos engineering to validate resilience assumptions.
- Security, Compliance & Governance: Enforce least-privilege IAM and centralised secrets management, maintain patch and vulnerability hygiene across hosts and images, and support audit requirements (SOC 2 / ISO 27001 / PCI-DSS) with CIS-aligned hardened baselines.
Must-Have Skills:
- AWS (Primary): EC2, VPC & network design, IAM, S3, RDS/Aurora, ELB/ALB, Route 53, Lambda, CloudWatch, CloudTrail, AWS Backup, Organizations / multi-account governance.
- Kubernetes / EKS: Production EKS ownership cluster provisioning & upgrades, Karpenter / Cluster Autoscaler, HPA, Helm, ingress, IRSA, RBAC, network policies, deep troubleshooting (CNI, CoreDNS, CSI).
- Azure (Secondary): Azure Backup & Recovery Services vaults (mandatory), Azure Site Recovery, VMs, VNets, Entra ID, Storage, AKS, Azure Monitor / Log Analytics.
- IaC: Terraform at expert level (modules, remote state, workspaces, drift detection, CI-driven plan/apply) and strong AWS CloudFormation (nested stacks, StackSets, change sets).
- Observability: Prometheus, Grafana, CloudWatch, Azure Monitor, ELK/OpenSearch, OpenTelemetry; plus one APM (Datadog / New Relic / Dynatrace). SLO and error-budget dashboards.
- SRE Practice: Hands-on SLI/SLO definition, error-budget policy, incident command, blameless postmortems, toil reduction, capacity planning, chaos/DR testing.
- CI/CD: GitHub Actions, GitLab CI, Jenkins or Azure DevOps; ArgoCD or Flux; blue-green and canary deployment with automated rollback.
Qualifications:
- Experience: 812 years in IT infrastructure, cloud or DevOps, including a minimum of 5 years in a dedicated SRE / DevOps / Cloud Infrastructure engineering role with production ownership.
- Education: Bachelor's degree in Computer Science, Engineering or equivalent practical experience.
- Certifications (preferred): AWS Solutions Architect / DevOps Engineer Professional; CKA or CKS; Azure AZ-104 or AZ-305; HashiCorp Terraform Associate.
- Soft skills: Strong ownership, calm and structured under production pressure, data-driven on reliability trade-offs, and able to explain risk clearly to non-technical stakeholders.
Good to Have:
- Go for operational tooling or Kubernetes operators; AWS CDK or Pulumi; service mesh (Istio / Linkerd / App Mesh) in production.
- Chaos engineering tooling (AWS FIS, Chaos Mesh, Gremlin); database reliability engineering (Aurora, PostgreSQL, DynamoDB); Kafka / MSK / Kinesis at scale.
- FinOps and cloud cost ownership; data-centre-to-cloud or AWS-to-Azure migration experience; regulated-industry exposure.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1665471