Posted on: 07/05/2026


Skills Required :
- Strong hands-on understanding of SRE principles including reliability engineering, observability, scalability, and automation.
- Proven experience with monitoring and logging platforms such as Dynatrace, Splunk, Prometheus, Grafana, and similar tools.
- Solid understanding of incident management, problem management, and change management with real-world ITIL application experience.
- Strong troubleshooting skills across operating systems, networking, applications, databases, and cloud infrastructure layers.
- Hands-on experience working with Kubernetes and OpenShift production environments.
- Good understanding of CI/CD pipelines, release processes, and production deployment activities.
Job Description :
- Improve operational hygiene by continuously enhancing alerts, dashboards, documentation, and on-call readiness processes.
- Perform incident triage and Root Cause Analysis (RCA) for recurring and complex production issues.
- Drive preventive and corrective actions based on incident trends, reliability insights, and operational
analysis.
- Update and maintain runbooks, SOPs, SLOs, SLIs, and error budgets.
- Identify opportunities to improve MTTA, MTTR, system availability, reliability, and scalability.
- Collaborate with Development teams to ensure reliability considerations are embedded into application design and release cycles.
- Support deployment activities and post-release validation processes to ensure production stability.
- Drive operational excellence initiatives by improving alerts, dashboards, documentation, and operational
processes.
- Monitor production systems proactively and ensure quick resolution of critical incidents.
- Work closely with cross-functional teams for troubleshooting and issue resolution.
- Mentor junior engineers and guide them during incidents, troubleshooting activities, and operational
support tasks.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1634253