Posted on: 01/09/2026
Responsibilities :
- Monitor, maintain, and improve production system availability, reliability, and performance.
- Design and manage Infrastructure as Code using Terraform or Ansible.
- Automate operational processes to improve system reliability and efficiency.
- Build and manage monitoring and alerting using Prometheus, Grafana, or Datadog.
- Troubleshoot production issues, perform RCA, and drive preventive fixes.
- Participate in on-call and incident management.
- Implement High Availability and Disaster Recovery strategies.
- Optimise application and infrastructure performance and address bottlenecks.
- Manage Kubernetes environments and workloads.
- Partner with engineering teams to implement SRE best practices.
Requirements :
- 3 - 5 years of experience in SRE, DevOps, or System Administration.
- Hands-on experience with AWS, GCP, or Azure.
- Strong experience with Docker and Kubernetes.
- Proficiency in Terraform or Ansible.
- Good programming skills in Python, Golang, or similar languages.
- Experience with CI/CD tools such as Jenkins or GitLab CI/CD.
- Exposure to monitoring and logging tools such as Prometheus, Grafana, Splunk, or Datadog.
- Good understanding of networking, load balancing, and caching.
- Bachelor's degree in Computer Science, Engineering, or equivalent experience.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1667705