HamburgerMenu
hirist

Site Reliability Engineer - Monitoring Tools

Build Up Rise Up UG
5 - 7 Years
Multiple Locations

Posted on: 09/07/2026

Job Description

Role : SRE Observability Engineer

We are currently hiring for an SRE Observability Engineer for a fast-growing software company in Pune, India.

Key Responsibilities :

- Design and implement end-to-end observability solutions using modern monitoring and logging platforms.

- Define and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Key Performance Indicators (KPIs) to ensure system reliability.

- Perform Root Cause Analysis (RCA), trend analysis, and support Problem Management initiatives.

- Build dashboards, alerting rules, capacity reports, and anomaly detection mechanisms.

- Analyze logs, metrics, and distributed traces to identify performance bottlenecks and improve system reliability.

- Identify monitoring gaps and enhance observability coverage across cloud-native applications.

- Support load and performance testing activities.

- Participate in Major Incident Management and ensure timely service recovery.

- Collaborate with DevOps and Engineering teams to improve system resilience, scalability, and operational excellence.

- Develop and validate recovery plans while driving continuous reliability improvements.

Required Skills / Primary Skills :

- Strong experience with Dynatrace, Grafana, Kibana, and the ELK Stack.

- Hands-on experience with Splunk and OpenTelemetry.

- Experience with cloud monitoring on AWS and/or Azure.

- Strong understanding of distributed tracing and observability best practices.

- Experience monitoring Kubernetes-based applications and infrastructure.

- Knowledge of Site Reliability Engineering (SRE) principles.

- Experience with monitoring, alerting, incident management, and performance optimization.

- Strong analytical, troubleshooting, and Root Cause Analysis (RCA) skills.

Additional Skills :

- SRE or DevOps certifications are preferred.

- Experience supporting high-volume, real-time distributed systems.

- Knowledge of cloud-native architectures and microservices.

- Excellent problem-solving and communication skills.

- Ability to collaborate effectively with cross-functional engineering and operations teams.

- Self-driven with a proactive approach to improving system reliability and performance.

Notice Period :

- Immediate Joiners Preferred.F

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...