Posted on: 04/07/2026
Job Description :
We are seeking a visionary Vice President SRE & AI Ops to lead the transformation of enterprise IT operations through AI-driven automation, Site Reliability Engineering (SRE), and Autonomous Operations. This leadership role will be responsible for building a self-healing, predictive, and highly automated IT ecosystem by leveraging AIOps, Agentic AI, and cloud-native technologies.
The ideal candidate will have a proven track record in leading large-scale infrastructure, SRE, DevOps, and IT operations teams while driving operational excellence, innovation, and digital transformation.
Key Responsibilities :
Strategic Leadership :
- Define and execute the enterprise-wide roadmap for Autonomous IT Operations, AIOps, and Agentic AI.
- Develop and implement strategies to build self-healing, predictive, and intelligent IT operations.
- Partner with business and technology leaders to align operational transformation initiatives with organizational goals.
- Establish governance frameworks, operational KPIs, and technology standards for enterprise reliability.
Operational Transformation :
- Lead the implementation of AI-powered monitoring, anomaly detection, predictive analytics, and intelligent remediation.
- Drive closed-loop automation across infrastructure, applications, and operational workflows.
- Build autonomous incident management capabilities through AI-driven root cause analysis and automated resolution.
- Achieve strategic operational goals, including:
1. Elimination of Level 1 (L1) operational support through automation.
2. Reduction of Level 2 (L2) support effort by at least 50% through AI-assisted operations.
Site Reliability Engineering & Infrastructure :
- Lead enterprise SRE practices focused on system reliability, scalability, resilience, and availability.
- Oversee hybrid cloud infrastructure across AWS, Azure, OCI, and on-premises environments.
- Manage enterprise Kubernetes platforms and cloud-native infrastructure.
- Define and govern Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, and reliability metrics.
- Drive disaster recovery (DR), business continuity, and data center modernization initiatives.
DevOps, Automation & AI Innovation :
- Modernize CI/CD pipelines by embedding AI-driven quality, deployment, and operational intelligence.
- Drive Infrastructure as Code (IaC), configuration management, and DevSecOps best practices.
- Establish a Center of Excellence (CoE) for AIOps, automation, and autonomous operations.
- Evaluate emerging AI technologies and implement innovative solutions to improve operational efficiency.
Incident & Service Management :
- Transform IT Service Management (ITSM) processes using AI-driven automation and self-service capabilities.
- Automate incident response, problem management, root cause analysis, and change management workflows.
- Leverage platforms such as ServiceNow to improve operational efficiency and service delivery.
- Ensure proactive monitoring, observability, and continuous service improvement.
Leadership & Stakeholder Management :
- Lead and mentor high-performing SRE, DevOps, Infrastructure, and Automation teams.
- Build a culture of reliability, innovation, accountability, and continuous improvement.
- Collaborate with executive leadership, business stakeholders, technology partners, and vendors.
- Present operational health, transformation progress, and strategic initiatives to senior leadership.
Required Skills :
- 15+ years of experience in Infrastructure Engineering, Site Reliability Engineering (SRE), DevOps, Cloud Operations, or IT Operations.
- Proven leadership experience in enterprise-scale infrastructure and operations transformation.
- Strong expertise in AIOps platforms and AI-driven operational automation.
- Hands-on experience with Agentic AI frameworks and autonomous operations.
- Deep knowledge of AWS, Microsoft Azure, Oracle Cloud Infrastructure (OCI), and hybrid cloud environments.
- Strong expertise in Kubernetes, container orchestration, and cloud-native technologies.
- Experience implementing observability platforms, monitoring solutions, distributed tracing, and logging frameworks.
- Strong understanding of SLOs, SLIs, error budgets, and reliability engineering principles.
- Expertise in Infrastructure as Code (Terraform, Ansible, CloudFormation), CI/CD, and DevSecOps.
- Experience with ServiceNow, ITSM automation, and AI-powered incident management.
- Strong knowledge of automation scripting using Python, Shell, or similar technologies.
- Excellent strategic planning, stakeholder management, and executive communication skills.
Preferred Qualifications :
- Bachelor's or Master's degree in Computer Science, Engineering, Information Technology, or a related field.
- Cloud certifications (AWS, Azure, OCI) and Kubernetes certifications are highly desirable.
- Certifications in ITIL, SRE, DevOps, or AI/ML technologies are an added advantage.
- Experience leading enterprise digital transformation and modernization initiatives.
- Exposure to Generative AI, Large Language Models (LLMs), and enterprise AI platforms is preferred.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Site Reliability Engineering
Job Code
1651361