Posted on: 30/05/2026
About the team :
As a Site Reliability Engineer 3, youll be part of the Platform & Infrastructure Engineering team.
The team is responsible for designing, automating, building, and operating highly reliable cloud-native systems.
About the role :
In this role you will be a core owner of production infrastructure and reliability for high-scale, regulated systems.
You will operate at the intersection of cloud infrastructure, Kubernetes platforms, security, observability, and incident management, enabling engineering teams to ship confidently while meeting stringent availability, scalability, and compliance requirements.
This role expects strong hands-on ownership, sound architectural judgment, and the ability to lead reliability initiatives end-to-end, not just execute tasks.
What you will do :
- Design, build, and operate AWS-based production infrastructure spanning networking, compute, storage, security, and observability.
- Own Kubernetes (EKS) platforms and critical production workloads at scale, including upgrades, capacity planning, and operational stability.
- Architect and manage VPCs, routing, security groups, NACLs, ALB/NLB, and hybrid or private connectivity patterns.
- Implement and operate service mesh (Istio) for traffic control, security policies, and service-level observability.
- Design fault-tolerant, highly available architectures across multiple AZs and regions.
- Define, implement, and continuously improve Disaster Recovery (DR) strategies with clearly articulated RPO/RTO aligned to business SLAs.
- Lead production readiness reviews, capacity planning, and reliability improvements for critical services.
- Build and standardize infrastructure automation using Terraform, including reusable modules, guardrails, and a clean Terraform SDLC.
- Enable GitOps-driven deployments using ArgoCD and CI/CD pipelines (GitHub Actions or equivalent).
- Reduce operational toil through automation and by building self-service platform capabilities.
- Build scalable metrics, alerting, and logging pipelines that support high-traffic, low-latency systems.
- Lead incident response, drive blameless post-mortems, and translate learnings into systemic fixes.
- Partner closely with Security teams to implement defense-in-depth using Cloudflare, network firewalls, IAM, and AWS security primitives.
- Deep dive into Linux, networking, and system performance issues under real production load.
- Mentor junior engineers, set technical standards, and raise the overall reliability bar for the Infra team.
What you will need :
- 9+ years of experience in SRE / DevOps / Platform / Infrastructure Engineering roles.
- Strong hands-on experience with AWS (EC2, ASG, ALB/NLB, IAM, VPC, networking).
- Proven experience operating Kubernetes (EKS) in production, high-availability environments.
- Deep understanding of networking fundamentals (routing, DNS, load balancing, firewalls).
- Strong experience with Terraform and infrastructure automation at scale.
- Excellent Linux fundamentals and production troubleshooting skills.
- Solid understanding of monitoring, alerting, and logging systems.
- Experience with Golang (preferred) or strong proficiency in Python / TypeScript.
- Prior exposure to service mesh architectures (Istio or similar).
- Experience operating systems with strict uptime, security, and compliance requirements.
- Experience leading incidents and reliability initiatives, not just participating in them.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1640467