Posted on: 26/08/2026
Job Description:
Serve as the final point of escalation for complex incidents.
- Develop advanced automation scripts and tools for deployment, monitoring, and recovery.
- Maintain and optimize CI/CD pipelines for zero-downtime deployments.
- Conduct system capacity planning and performance tuning.
- Analyze metrics and logs to identify and resolve bottlenecks.
- Collaborate with development teams to embed SRE principles into the product lifecycle.
Required Skills & Qualifications:
- 8-12 years of experience in SRE/DevOps.
- Expertise in: Cloud Platforms: AWS; Containerization: Kubernetes.
- Infrastructure as Code: Terraform, Ansible.
- Proficiency with: Monitoring tools - Prometheus, and/or Grafana.
- Strong with CI/CD tools - Jenkins, AWS Code Deployment, GitLab CI, and ArgoCD.
- Database management (SQL/NoSQL) and caching strategies.
- Deep knowledge of networking, security, and Linux/Windows systems.
- Strong leadership, communication, and problem-solving skills.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1666172