HamburgerMenu
hirist

Observability & Monitoring Lead - DevOps

Tidyhire
12 - 16 Years
Hyderabad

Posted on: 05/08/2026

Job Description

Key Responsibilities:

- Lead the design, implementation, and governance of enterprise observability and monitoring solutions.

- Build and maintain monitoring platforms using Prometheus, Grafana, Azure Monitor, Application Insights, Splunk, and OpenTelemetry.

- Develop dashboards, alerts, and SLO/SLI metrics to proactively monitor application and infrastructure health.

- Monitor and optimize the performance, availability, and reliability of cloud platforms, Kubernetes clusters, and business-critical applications.

- Lead major incident response, root cause analysis (RCA), and continuous service improvement initiatives.

- Implement centralized logging, distributed tracing, and end-to-end observability across cloud-native environments.

- Automate monitoring deployment, alerting, and operational workflows using Terraform, Ansible, Python, PowerShell, or Bash.

- Collaborate with Platform Engineering, DevOps, SRE, Security, and Application teams to establish monitoring standards and best practices.

- Drive capacity planning, performance tuning, and cost optimization for monitoring platforms.

- Mentor engineers on observability practices and foster a culture of reliability, automation, and operational excellence.

Key Skills and Experience:

- 1216 years of overall IT experience, with at least 6+ years of hands-on experience in Observability, Monitoring, Site Reliability Engineering (SRE), or Platform Engineering.

- Proven experience in designing, implementing, and managing enterprise observability and monitoring solutions for large-scale production environments.

- Strong hands-on expertise with Prometheus, Grafana, Azure Monitor, Application Insights, OpenTelemetry, and centralized logging platforms such as Splunk or ELK.

- Extensive experience monitoring Azure cloud infrastructure, Kubernetes (AKS), containerized applications, distributed systems, and cloud-native platforms.

- Hands-on experience in building dashboards, alerts, metrics, logs, and distributed tracing to ensure end-to-end visibility and proactive monitoring.

- Strong understanding of Site Reliability Engineering (SRE) principles, including SLIs, SLOs, Error Budgets, Incident Management, Root Cause Analysis (RCA), performance optimization, and service reliability.

- Experience supporting enterprise production environments, managing major incidents, and driving high availability, operational excellence, and continuous service improvement.

- Proficiency in automation and Infrastructure as Code (IaC) using Terraform, Ansible, Python, PowerShell, or Bash.

- Strong troubleshooting, analytical, communication, stakeholder management, and cross-functional collaboration skills.

Preferred Certifications:

- Microsoft Certified: Azure Administrator (AZ-104)

- Microsoft Certified: Azure Solutions Architect (AZ-305)

- Certified Kubernetes Administrator (CKA)

- Grafana Certified Professional

- Prometheus Certified Associate (PCA)

- ITIL Foundation

- SRE Foundation / SRE Practitioner

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...