Posted on: 06/05/2026
Required Skills & Expertise :
Core SRE & Reliability :
- Strong understanding of SRE principles and practices (SLIs, SLOs, SLAs, Error Budgets).
- Proven experience in SRE transition or transformation programs.
- Expertise in high availability, scalability, fault tolerance, and resiliency engineering.
- Strong experience with incident management, escalation, and post-incident reviews.
DevOps & Automation :
- Hands-on experience with CI/CD pipelines (Jenkins, GitLab CI, Azure DevOps, or similar).
- Strong scripting/programming skills in Python, Go, Shell, or Java.
- Experience with Infrastructure as Code (IaC) tools such as Terraform, CloudFormation, or ARM.
Monitoring & Observability :
- Expertise in monitoring and observability tools such as Prometheus, Grafana, ELK/EFK, Datadog, New Relic, Splunk.
- Experience defining alerts that align with SLOs and reduce alert fatigue.
Cloud & Platforms :
- Strong experience with cloud platforms (AWS, Azure, or Google Cloud).
- Hands-on experience with containerization and orchestration (Docker, Kubernetes).
- Experience managing distributed systems and microservices architectures.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1633780