HamburgerMenu
hirist

Senior Observability Engineer - Grafana/Prometheus

ApTask
6 - 10 Years
Multiple Locations

Posted on: 01/09/2026

Job Description

Job Title : Senior Observability Engineer / Platform Engineer

Experience : 6 - 10 years

Job Location : Remote (Hyderabad, Blore, Delhi-NCR, Mumbai, Pune, Bhuvaneswar, Coimbatore, Jaipur etc) - one virtual round, 1 F2F meeting (mandatory), 1 client round of technical interviews.

Timings : 12 pm to 10 pm

Job Description :

We are looking for a hands-on Observability Engineer with strong experience in cloud-native platforms, monitoring, logging, alerting, and operational excellence. The ideal candidate will have experience building and managing enterprise observability solutions across Kubernetes and public cloud environments, enabling proactive monitoring, incident response, and reliability engineering practices.

Key Responsibilities :

- Design, implement, and maintain enterprise observability platforms covering metrics, logs, traces, and events.

- Build observability solutions using tools such as Prometheus, Grafana, OpenSearch, Splunk, Elastic, Datadog, Dynatrace, New Relic, or equivalent platforms.

- Develop dashboards, SLOs, SLIs, alerting rules, and service health monitoring frameworks.

- Integrate monitoring and observability capabilities within Kubernetes and cloud-native environments.

- Enable incident management, root cause analysis, and operational troubleshooting through observability best practices.

- Automate monitoring configuration and platform onboarding using Infrastructure-as-Code and CI/CD pipelines.

- Collaborate with engineering, platform, SRE, and operations teams to improve reliability, performance, and availability.

- Support observability maturity initiatives including distributed tracing, AIOps, and intelligent alerting.

Required Skills :

- 6+ years of experience in Platform Engineering, SRE, DevOps, Cloud Operations, or Observability Engineering.

- Strong expertise in : Prometheus, Grafana, Splunk / OpenSearch / Elastic, Distributed tracing solutions (Jaeger, Tempo, Open Telemetry, etc.).

- Good understanding of Kubernetes, Docker, and containerized workloads.

- Experience with AWS, Azure, or GCP environments. Knowledge of incident management, alert tuning, and troubleshooting production environments.

- Experience with scripting and automation using Python, Shell, or similar languages.

- Familiarity with CI/CD and Infrastructure-as-Code tools such as Terraform, Jenkins, GitHub Actions, or Argo CD.

- Preferred Skills : Exposure to SRE practices, error budgets, SLO/SLI frameworks.

- Experience with AIOps, automated remediation, or intelligent incident response. Knowledge of Open Telemetry implementation.

- Working experience in enterprise-scale production platforms. Exposure to security and governance considerations in cloud-native environments.

The job is for:

May work from home
info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...