Posted on: 09/07/2026
Role : SRE Observability Engineer
We are currently hiring for an SRE Observability Engineer for a fast-growing software company in Pune, India.
Key Responsibilities :
- Design and implement end-to-end observability solutions using modern monitoring and logging platforms.
- Define and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Key Performance Indicators (KPIs) to ensure system reliability.
- Perform Root Cause Analysis (RCA), trend analysis, and support Problem Management initiatives.
- Build dashboards, alerting rules, capacity reports, and anomaly detection mechanisms.
- Analyze logs, metrics, and distributed traces to identify performance bottlenecks and improve system reliability.
- Identify monitoring gaps and enhance observability coverage across cloud-native applications.
- Support load and performance testing activities.
- Participate in Major Incident Management and ensure timely service recovery.
- Collaborate with DevOps and Engineering teams to improve system resilience, scalability, and operational excellence.
- Develop and validate recovery plans while driving continuous reliability improvements.
Required Skills / Primary Skills :
- Strong experience with Dynatrace, Grafana, Kibana, and the ELK Stack.
- Hands-on experience with Splunk and OpenTelemetry.
- Experience with cloud monitoring on AWS and/or Azure.
- Strong understanding of distributed tracing and observability best practices.
- Experience monitoring Kubernetes-based applications and infrastructure.
- Knowledge of Site Reliability Engineering (SRE) principles.
- Experience with monitoring, alerting, incident management, and performance optimization.
- Strong analytical, troubleshooting, and Root Cause Analysis (RCA) skills.
Additional Skills :
- SRE or DevOps certifications are preferred.
- Experience supporting high-volume, real-time distributed systems.
- Knowledge of cloud-native architectures and microservices.
- Excellent problem-solving and communication skills.
- Ability to collaborate effectively with cross-functional engineering and operations teams.
- Self-driven with a proactive approach to improving system reliability and performance.
Notice Period :
- Immediate Joiners Preferred.F
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1652828