Posted on: 20/08/2026
About the Role :
We are looking for a Senior Site Reliability Engineer to own the reliability of our production platform end-to end. We run a Kubernetes-native, multi-region platform (Singapore, Hong Kong, Sydney) serving a regulated digital wealth-management product, where availability, latency, and trust are first-order product features.
This is a senior individual-contributor role, not a managerial one. You will define what "reliable" means in measurable terms, build the systems and automation that keep us there, and own and lead our on-call and incident-response program.
You'll spend your time engineering reliability into the platform through SLOs, observability, and automation rather than firefighting, and you'll raise the bar for how the whole engineering org operates production.
What You'll Own :
- Reliability targets : Define and drive SLIs/SLOs and error budgets across critical services; partner with product and engineering teams to make error-budget-based decisions that balance velocity and stability.
- On-call & incident program : Own the on-call rotation, escalation policies, and paging strategy. Establish incident command, run blameless postmortems, and turn RCAs into tracked, completed reliability work. Drive down MTTD and MTTR.
- Platform & Kubernetes reliability. Own the reliability of our EKS-based deployment platform GitOps delivery (ArgoCD), Helm-based release configuration, and Infrastructure as Code (Terraform/OpenTofu) on AWS. Make deployments safe, progressive, and reversible.
Qualifications (Must-have) :
- Bachelor's or Master's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 4 - 8 years in SRE, platform engineering, or DevOps, with a strong senior IC track record of owning production systems.
- Production-grade expertise with Kubernetes and containers, and a cloud platform (AWS preferred) in a distributed-systems environment.
- Hands-on experience defining and operating against SLIs/SLOs and error budgets, and leading incident response and blameless postmortems.
- Strong with observability tooling metrics, logging, tracing, dashboards, and alerting (e.g., Datadog, Prometheus/Grafana, or equivalents).
- Proficient writing automation and infrastructure code (e.g., Python, Go, Shell, Terraform) and comfortable with GitOps / CI/CD delivery.
Nice to Have :
- Experience operating in regulated or high-trust environments (fintech, payments, etc.).
- Experience running multi-region / multi-cluster Kubernetes at scale.
- Familiarity with ArgoCD, Helm/Helmfile, OpenTofu, Vault, Cloudflare, or comparable tooling.
- Experience building self-service developer platforms or internal reliability tooling.
- Contributions to open-source projects or public technical content (GitHub, blogs, talks).
- Relevant cloud or Kubernetes certifications (e.g., AWS, CKA/CKS).
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1664601