HamburgerMenu
hirist

Vice President - Site Reliability & AIOps

Talent Toppers
14 - 18 Years
Gurgaon/Gurugram

Posted on: 04/07/2026

Job Description

Job Description :

We are seeking a visionary Vice President SRE & AI Ops to lead the transformation of enterprise IT operations through AI-driven automation, Site Reliability Engineering (SRE), and Autonomous Operations. This leadership role will be responsible for building a self-healing, predictive, and highly automated IT ecosystem by leveraging AIOps, Agentic AI, and cloud-native technologies.

The ideal candidate will have a proven track record in leading large-scale infrastructure, SRE, DevOps, and IT operations teams while driving operational excellence, innovation, and digital transformation.

Key Responsibilities :

Strategic Leadership :

- Define and execute the enterprise-wide roadmap for Autonomous IT Operations, AIOps, and Agentic AI.

- Develop and implement strategies to build self-healing, predictive, and intelligent IT operations.

- Partner with business and technology leaders to align operational transformation initiatives with organizational goals.

- Establish governance frameworks, operational KPIs, and technology standards for enterprise reliability.

Operational Transformation :

- Lead the implementation of AI-powered monitoring, anomaly detection, predictive analytics, and intelligent remediation.

- Drive closed-loop automation across infrastructure, applications, and operational workflows.

- Build autonomous incident management capabilities through AI-driven root cause analysis and automated resolution.

- Achieve strategic operational goals, including:

1. Elimination of Level 1 (L1) operational support through automation.

2. Reduction of Level 2 (L2) support effort by at least 50% through AI-assisted operations.

Site Reliability Engineering & Infrastructure :

- Lead enterprise SRE practices focused on system reliability, scalability, resilience, and availability.

- Oversee hybrid cloud infrastructure across AWS, Azure, OCI, and on-premises environments.

- Manage enterprise Kubernetes platforms and cloud-native infrastructure.

- Define and govern Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, and reliability metrics.

- Drive disaster recovery (DR), business continuity, and data center modernization initiatives.

DevOps, Automation & AI Innovation :

- Modernize CI/CD pipelines by embedding AI-driven quality, deployment, and operational intelligence.

- Drive Infrastructure as Code (IaC), configuration management, and DevSecOps best practices.

- Establish a Center of Excellence (CoE) for AIOps, automation, and autonomous operations.

- Evaluate emerging AI technologies and implement innovative solutions to improve operational efficiency.

Incident & Service Management :

- Transform IT Service Management (ITSM) processes using AI-driven automation and self-service capabilities.

- Automate incident response, problem management, root cause analysis, and change management workflows.

- Leverage platforms such as ServiceNow to improve operational efficiency and service delivery.

- Ensure proactive monitoring, observability, and continuous service improvement.

Leadership & Stakeholder Management :

- Lead and mentor high-performing SRE, DevOps, Infrastructure, and Automation teams.

- Build a culture of reliability, innovation, accountability, and continuous improvement.

- Collaborate with executive leadership, business stakeholders, technology partners, and vendors.

- Present operational health, transformation progress, and strategic initiatives to senior leadership.

Required Skills :

- 15+ years of experience in Infrastructure Engineering, Site Reliability Engineering (SRE), DevOps, Cloud Operations, or IT Operations.

- Proven leadership experience in enterprise-scale infrastructure and operations transformation.

- Strong expertise in AIOps platforms and AI-driven operational automation.

- Hands-on experience with Agentic AI frameworks and autonomous operations.

- Deep knowledge of AWS, Microsoft Azure, Oracle Cloud Infrastructure (OCI), and hybrid cloud environments.

- Strong expertise in Kubernetes, container orchestration, and cloud-native technologies.

- Experience implementing observability platforms, monitoring solutions, distributed tracing, and logging frameworks.

- Strong understanding of SLOs, SLIs, error budgets, and reliability engineering principles.

- Expertise in Infrastructure as Code (Terraform, Ansible, CloudFormation), CI/CD, and DevSecOps.

- Experience with ServiceNow, ITSM automation, and AI-powered incident management.

- Strong knowledge of automation scripting using Python, Shell, or similar technologies.

- Excellent strategic planning, stakeholder management, and executive communication skills.

Preferred Qualifications :

- Bachelor's or Master's degree in Computer Science, Engineering, Information Technology, or a related field.

- Cloud certifications (AWS, Azure, OCI) and Kubernetes certifications are highly desirable.

- Certifications in ITIL, SRE, DevOps, or AI/ML technologies are an added advantage.

- Experience leading enterprise digital transformation and modernization initiatives.

- Exposure to Generative AI, Large Language Models (LLMs), and enterprise AI platforms is preferred.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...