Posted on: 30/05/2026
About the role :
Were looking for a Senior Staff Engineer (Infrastructure) who can drive initiatives that improve system reliability, accelerate release cycles, and build scalable infrastructure systems.
This role is a mix of hands-on engineering, technical leadership, and strategy, with a strong emphasis on building tools and frameworks that directly impact developer velocity and operational excellence.
What You Will Do :
- Lead the Site Reliability / Infrastructure Engineering team, setting the technical vision and execution roadmap for reliability, scalability, and availability.
- Design and operate highly available, fault-tolerant infrastructure systems that support mission-critical services.
- Own and evolve CI/CD, release, and deployment reliability with strong focus on stability, rollback, and blast-radius reduction.
- Drive adoption of cloud-native and distributed systems practices (Kubernetes, AWS, multi-region architectures, service monitoring and alerting).
- Partner with engineering leaders to define and enforce SLOs, SLIs, error budgets, and reliability standards across the organization.
- Champion best practices in incident management, on-call excellence, postmortems, and continuous improvement.
- Lead capacity planning, performance engineering, and cost optimization initiatives at scale.
- Mentor senior engineers, build strong on-call and operations culture, and foster engineering rigor and ownership.
What You Will Need :
- 10+ years in software engineering with at least 3 years in hands on site reliability roles.
- Strong expertise in Kubernetes, AWS, CI/CD pipelines, Spinnaker/Jenkins/ArgoCD, IaC.
- Hands-on experience with test automation frameworks, dynamic test selection, code quality gates.
- Strong coding skills in Go, Python, or Java (Go preferred for tooling/frameworks).
- Experience designing ephemeral environments, monitoring/observability tools.
- Proven track record of leading small teams, mentoring engineers, and driving large-scale initiatives.
- A background in Google Summer of Code (GSoC) or any open-source software (OSS) contributions would be a huge plus.
- Alternatively, having built something really cool to solve a difficult pain point for engineers or anything that helped ship things faster for the company or the OSS community would also be great.
- Strong communication and ability to influence cross-functional stakeholders.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1640465