Posted on: 18/07/2026
Job Description :
Responsibilities :
- Stack Management : Deploy, configure, and maintain the core observability stack using Prometheus, Grafana, Alertmanager and Loki.
- Dashboarding & Visualization : Collaborate with engineering teams to design and build comprehensive Grafana dashboards for application and infrastructure health monitoring.
- Alerting Strategy : Configure and fine-tune Alertmanager rules to ensure accurate, actionable alerts while minimizing alert fatigue.
- Log Management : Architect and manage centralized logging solutions using Loki to ensure efficient log aggregation and querying.
- System Optimization : Monitor the performance of the observability stack itself, optimizing resource usage and scaling infrastructure as needed.
- Continuous Improvement : Evaluate and integrate modern observability tools (such as VictoriaMetrics and VictoriaLogs) to enhance system correlation and analysis capabilities.
- Experience : 9+ years of hands-on experience in DevOps, SRE or Observability roles.
- Core Stack : Deep technical expertise in Prometheus, Grafana, Alertmanager, and Loki (PLG Stack).
- Automation : Proficiency in scripting (Python, Bash) and Infrastructure as Code (e.g., Terraform, Ansible).
- Infrastructure & OS : Strong working knowledge of Linux/Unix administration.
- Containerization : Experience monitoring containerized environments (Docker, Kubernetes).
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
DevOps / Cloud
Job Code
1655565