Posted on: 18/09/2026
Role & responsibilities :
The Senior Cloud Site Reliability Engineer (Senior Cloud SRE) is responsible for ensuring the reliability, scalability, availability, performance, security, and operational excellence for cloud platforms and critical product infrastructure.
This role combines software engineering, cloud engineering, automation, observability, and operational governance practices to build highly resilient and self-healing platforms across hybrid and cloud-native environments.
Site Reliability Engineering & Operational Excellence :
- Drive and implement Site Reliability Engineering (SRE) best practices across cloud platforms and services.
- Define, maintain, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), Service Level Agreements (SLAs), and Error Budgets.
- Improve service reliability, resiliency, scalability, and operational efficiency.
- Establish operational standards, reliability governance, and production readiness practices.
- Conduct Root Cause Analysis (RCA), postmortems, and reliability improvement initiatives.
- Participate in on-call rotations, incident management, and major incident resolution activities.
- Continuously improve operational processes to reduce Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR).
Observability, Monitoring & Telemetry :
- Design, implement, and maintain enterprise observability and telemetry platforms.
- Build operational dashboards, reliability scorecards, and service health monitoring solutions.
- Configure proactive alerting, anomaly detection, and incident correlation mechanisms.
- Implement centralized monitoring and telemetry using Grafana, Prometheus, Azure Monitor, Log Analytics, ELK Stack / ElasticSearch, and Power BI dashboards.
Automation & Auto-Healing Engineering :
- Drive automation-first operational practices across infrastructure and platform services.
- Develop Infrastructure-as-Code (IaC) solutions using Terraform, ARM/Bicep, and Ansible.
- Build operational automation scripts using Python, Bash, and PowerShell.
- Develop self-healing and auto-remediation capabilities for recurring operational incidents.
Requirements :
- Bachelors degree in Computer Science, Engineering, or related field.
- Experience managing cloud platforms (Azure, AKS, or equivalent).
- Strong understanding of Kubernetes and containerization concepts.
- Knowledge of Python, scripting, and Infrastructure-as-Code tools (Terraform, Ansible).
- Solid knowledge of relational databases (MS-SQL) and exposure to NoSQL technologies (Redis, MongoDB).
- Experience with CI/CD tools (Azure DevOps, Jenkins, GitHub Actions).
- Strong Linux administration and troubleshooting skills.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1672758