Posted on: 31/08/2026
What You Will Do :
- Monitor, maintain, and improve the availability and performance of production systems.
- Design and manage Infrastructure as Code (IaC) with tools like Terraform or Ansible.
- Automate operational workflows to enhance system reliability.
- Build and maintain alerting and monitoring systems using Prometheus, Grafana, or Datadog.
- Perform root cause analysis and resolve recurring issues.
- Participate in on-call rotations for incident management.
- Implement disaster recovery and high availability strategies.
- Optimize performance bottlenecks in applications and systems.
- Manage Kubernetes-based environments and workloads.
- Collaborate with developers to implement SRE best practices.
What We Are Looking For :
- Bachelors degree in Computer Science, Engineering, or equivalent experience.
- 3-5 years of experience in SRE, DevOps, or system administration roles.
- Hands-on experience with cloud platforms (AWS, GCP, Azure).
- Experience with containerized applications (Docker, Kubernetes).
- Proficiency with automation tools (Terraform, Ansible).
- Strong programming skills in Python, Golang, or similar languages.
- Experience with CI/CD tools (Jenkins, GitLab CI/CD).
- Knowledge of monitoring/logging systems (Splunk, Datadog).
- Understanding of networking protocols, load balancing, and caching.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1667236