HamburgerMenu
hirist

Observability Engineer - Grafana/Prometheus

TGS The Global Skills
5 - 8 Years
Multiple Locations

Posted on: 25/05/2026

Job Description

Description :

We are specifically looking for a Grafana and Prometheus expert, not a traditional DevOps Engineer. This role is focused on Observability Engineering, Reliability Monitoring, and Deep Metrics Intelligence. The ideal candidate should have strong expertise in PromQL, advanced Grafana dashboarding, monitoring architecture, alerting strategies, and transforming raw logs/metrics into actionable operational insights. Candidates whose experience is primarily around CI/CD, Terraform setup, infrastructure provisioning, or standard cloud DevOps operations will not fit this requirement.

Responsibilities :

- Develop, optimize, and troubleshoot complex PromQL queries to extract actionable metrics from Prometheus.


- Design advanced Grafana dashboards with dynamic variables, transformations, drill-downs, and multi source integrations.

- Configure Prometheus Service Discovery for auto-scaling and dynamic target discovery.

- Build monitoring solutions for large-scale distributed systems and microservices.

- Implement proactive alerting strategies using Alertmanager and anomaly detection techniques.

- Integrate multiple observability data sources, including Prometheus, SQL, Elasticsearch, logs, and cloud metrics.

- Create correlated dashboards combining metrics, logs, and traces into a unified operational view.

- Support monitoring, incident response, root cause analysis, and reliability engineering initiatives.

- Build observability workflows that support self-healing systems and automated triggers.

- Collaborate directly with client stakeholders and offshore teams.

Requirements :

- We are specifically looking for candidates whose core expertise is Prometheus, PromQL, Grafana, Monitoring Architecture, Observability Engineering, Reliability Engineering.


- Strong PromQL projects.


- Experience in Advanced Grafana dashboards, Monitoring automation.

- Know about Alerting frameworks, Service Discovery implementations.

- Knowledge of Multi-source observability platforms, Incident response ownership.

- Experience in Reliability engineering contributions.

Must Have Skills :

- Expert-level Prometheus and PromQL experience (Mandatory).


- Advanced Grafana Dashboard Engineering.

- Strong understanding of Prometheus architecture, exporters, scraping configs, and recording rules.

- Experience with Alertmanager, Exporters, and monitoring ecosystems.

- Service Discovery configuration experience.

- Multi-source dashboard integration.

- AWS Cloud knowledge.

- Large-scale data and log management experience.

- Strong understanding of Monitoring and Incident Response workflows.

- Excellent English communication skills.

Strongly Preferred :

- Experience with Loki, Tempo, Flux, and Elasticsearch.


- Experience creating scalable dashboards for hundreds of microservices.

- Ability to design dashboards that tell operational stories through data visualization.

- Experience with anomaly detection and self-healing monitoring systems.

- Observability-first mindset focused on Reliability and Visibility.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...