HamburgerMenu
hirist

Infosys - Site Reliability Engineering Lead

EdgeVerve Systems
10 - 15 Years
Bangalore

Posted on: 20/08/2026

showcase-imageshowcase-image

Job Description

Role Overview :

- 10+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Operations.

- Deep hands-on experience with distributed systems, container orchestration (Kubernetes), and cloud-native operational tooling.

- Proficiency with automation and scripting languages (Python, Go, PowerShell, Ansible).

- Strong understanding of observability platforms (Splunk, Dynatrace) and event-driven monitoring.

- Proven leadership in major incident management and cross-team technical coordination.

- Strong grasp of networking, Linux/Unix internals, and modern infrastructure patterns.

- Excellent communication skills, including executive-level situational awareness during critical incidents.

- Demonstrated ability to influence technical roadmaps and drive adoption of reliability best practices.

Reliability Engineering & Automation :

- Architect and deliver automation solutions that eliminate toil, reduce MTTR, and increase service resilience.

- Experience in Ansible, Puppet or Chef is a plus.

- Implement intelligent alerting, anomaly detection, and event correlation leveraging AI and AIOps tools.

- Guide and enforce SLO/SLI adoption across product teams, ensuring metrics inform decision-making and prioritization.

- Utilize Infrastructure-as-Code (IaC) tools for automating deployment of assets within cloud tenants.

Observability & Operational Excellence :

- Ensure operational readiness of applications and platforms through resiliency testing, chaos engineering, and failure-mode validation.

Cross-Functional Leadership & Influence :

- Partner with Delivery, Architecture, Security, and Risk teams to embed reliability and resilience into design and execution.

Standardization & Documentation :

- Develop, maintain, and enforce runbooks, response playbooks, and automated recovery patterns.

- Follow best practices and internal processes for Non-Functional requirements to improve resiliency and reliability.

Mentorship & Technical Development :

- Coach and mentor Associate, Professional, and Senior SREs to build technical depth and operational discipline.

- Provide thought leadership in SRE methodologies, cloud-native operational patterns, and automated reliability engineering.

Incident Leadership & Production Operations :

- Lead P1/P0 incident bridges and direct technical investigation efforts.

- Perform hands-on triage using logs, traces, metrics, and application telemetry.

- Drive mitigation, recovery, RCA development, and follow-through remediation.

- Provide executive communications during major incidents.

- Build operational automation based on recurring production issues.

- Establish credibility through technical leadership during live service disruptions.

Additional Qualifications :

- Experience enabling large-scale SRE transformations or modernization initiatives.

- Demonstrated proficiency with GitLab Duo, or similar AI technologies.

- Familiarity with chaos engineering, resilience assessments, and service failure modeling.

- Exposure to hybrid-cloud and multi-cloud operational frameworks.

- Experience contributing to or leading Center for Enablement functions or Communities of Practice.

- Expertise with highly regulated industries preferred.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...