Posted on: 30/09/2026
Role Summary :
We are seeking a Senior Site Reliability Engineer (SRE) with deep expertise in observability, monitoring framework design, and Azure platform operations. This is a hands-on technical leadership role responsible for owning and driving the end-to-end monitoring framework for all Atlas5 production environments.
Key Responsibilities :
Monitoring Framework Ownership :
- Own the design, implementation, and ongoing management of the Atlas5 monitoring framework across all 5 layers : Infrastructure, Data Tier, Application, Integration, and File Exchange.
- Define and configure alert thresholds, severity levels, and escalation paths.
- Ensure every alert triggers a Jira incident ticket and DevOps channel notification within defined SLA.
- Enforce mandatory RCA completion for all P1-Critical incidents.
Tool Selection & Implementation :
- Evaluate and implement monitoring toolsets (Azure Monitor, Prometheus, Grafana, New Relic, SigNoz, Zabbix, Dynatrace).
- Build a cost-effective hybrid monitoring stack.
- Integrate monitoring tools with Jira Service Management.
- Build and maintain unified monitoring dashboards.
SRE Practice & Incident Management :
- Define and maintain SLOs, SLIs, and error budgets.
- Lead incident response for P1-Critical production incidents.
- Configure auto-escalation for unacknowledged Critical alerts.
- Conduct regular alert quality reviews.
Azure Observability Engineering :
- Configure and manage Azure Monitor, Log Analytics, and Application Insights.
- Design and manage KQL queries for alerting and reporting.
- Optimize monitoring costs through Metric Alert migration.
- Instrument applications with Open Telemetry.
Required Experience & Skills :
- 8 - 10 years of experience in SRE, observability engineering, or cloud operations.
- Proven experience designing monitoring frameworks from scratch in Azure.
- Strong Azure Monitor expertise (Log Analytics, Application Insights, KQL).
- Hands-on experience with at least 2 of : Prometheus, Grafana, New Relic, Datadog, Dynatrace, Zabbix, SigNoz.
- Experience integrating monitoring with incident management tools (Jira, PagerDuty).
- Scripting : KQL (mandatory), PowerShell or Python.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1675861