HamburgerMenu
hirist

AIOps Support Lead/Manager - Observability

Magnet HR Consulting Services
8 - 13 Years
Bangalore

Posted on: 25/09/2026

Job Description

We are seeking an experienced and dynamic AIOps Support Lead (Manager) to lead a modern, standardized observability and monitoring capability across a diverse enterprise application landscape.

The role will manage a 14-member Tier 1 AIOps Support team, initially supporting approximately 50 applications and scaling to nearly 320 applications over the next two years. The position will own people management, operational processes, triage quality, telemetry/data discipline, and continuous automation initiatives.

The key objective is to transition the organization from reactive, manual incident triage toward proactive and increasingly automated monitoring, supported by a mature OpenTelemetry-based observability framework.

Key Responsibilities :

1. Team & Shift Management :

- Lead, hire, coach, develop, and manage a team of 14 Tier 1 AIOps Support Engineers.

- Manage team performance, productivity, capability development, and resource allocation.

- Design and maintain effective 24x7 shift and roster coverage.

- Scale team operations in line with application portfolio growth.

- Establish clear roles, responsibilities, operating standards, and performance expectations.

2. Operations & Triage Management :

- Own the quality, consistency, and speed of incident detection, classification, and routing.

- Establish and monitor standards for Mean-Time-to-Triage (MTTT).

- Develop a structured and repeatable framework for onboarding applications into the AIOps monitoring ecosystem.

- Scale application onboarding from approximately 50 to 320 applications.

- Lead major outage and incident response activities.

- Coordinate with application, infrastructure, cloud, and other technical teams during critical incidents.

- Drive Root Cause Analysis (RCA) following major incidents.

- Develop and continuously improve technical runbooks, operating procedures, and knowledge repositories.

3. Observability & Data Quality :

- Establish standards for ingesting and managing heterogeneous telemetry and log data.

- Partner with application and technology teams to improve the quality, consistency, and usability of monitoring data.

- Support the transition toward a standardized OpenTelemetry backbone.

- Improve alert classification, correlation, enrichment, and prioritization.

- Reduce alert noise and false positives through continuous monitoring and optimization.

- Establish strong data-quality governance to ensure reliable automated triage.

4. Automation & Proactive Monitoring :

- Drive continuous reduction of manual triage activities through automation.

- Identify opportunities for automated alert correlation, event enrichment, incident routing, and early-warning detection.

- Support the transition from reactive monitoring to proactive and predictive operations.

- Prepare the observability environment for future AI-driven and Agentic AI automation.

- Collaborate with engineering and platform teams to identify automation use cases and improve operational efficiency.

5. Incident, Problem & Change Management :

- Ensure adherence to ITIL-based Incident, Problem, and Change Management processes.

- Drive effective incident escalation and resolution workflows.

- Ensure appropriate linkage between incidents, problems, changes, configuration items, and business services.

- Maintain alignment with CMDB and service-management practices.

- Monitor adherence to SLA/SLO commitments and operational governance standards.

6. KPI & Governance Reporting :

- Define and track operational KPIs and service-quality metrics.

- Prepare regular governance reports for program and technology leadership.

- Monitor metrics including :

1. SLA/SLO compliance

2. Mean-Time-to-Triage (MTTT)

3. Alert-to-incident accuracy

4. False-positive rate

5. Incident volume and trends

6. Automation percentage

7. Manual triage reduction

8. Application onboarding progress

- Present insights, trends, risks, and improvement opportunities to senior stakeholders.

Technical Requirements :

IT Service Management & Governance :

- Strong understanding of ITIL Incident Management lifecycle.

- Knowledge of Problem Management and Change Management.

- Understanding of CMDB concepts and service relationships.

- Strong understanding of SLA/SLO definitions and availability calculations.

- Experience working within structured operational governance frameworks.

Observability & Monitoring :

- Hands-on experience with one or more modern monitoring and observability platforms, such as Dynatrace, New Relic, AWS CloudWatch, ManageEngine, Glassbox, or similar enterprise monitoring/observability platforms.

- Additional exposure to OpenTelemetry, distributed tracing, Application Performance Monitoring (APM), log aggregation, metrics and event monitoring, and alert correlation and enrichment.

Infrastructure & Networking :

- Good understanding of Virtual Machines (VMs), firewalls, load balancers, containers, Kubernetes, OpenShift / OCP, and enterprise infrastructure and application architecture.

Scripting, APIs & Data Formats :

- Working knowledge of Unix/Linux shell scripting and exposure to Windows batch scripting.

- Understanding of JSON and XML data formats.

- Experience with API testing and troubleshooting tools such as Postman, SOAP UI, or similar API testing tools.

Cloud & Security :

- Understanding of cloud fundamentals including compute, storage, networking, and cloud monitoring.

- Basic understanding of security technologies and protocols including TLS, SSL, authentication tokens, and secrets management.

- Awareness of data infrastructure concepts such as data lineage, data latency, and data quality.

Leadership & Behavioral Competencies :

Team Leadership :

- Proven experience managing and scaling technical support or operations teams.

- Strong people-management, coaching, mentoring, and performance-management skills.

- Ability to manage teams operating in 24x7 support environments.

Scale-Up & Transformation :

- Demonstrated experience scaling support operations during periods of significant application or business growth.

- Ability to establish standardized processes while maintaining operational flexibility.

Analytical & Problem-Solving Skills :

- Ability to analyze fragmented and high-volume alert data.

- Strong capability to convert noisy monitoring signals into actionable insights.

- Ability to develop effective technical runbooks and operational procedures.

Adaptability :

- Comfortable operating in complex and ambiguous environments.

- Ability to dynamically prioritize applications and support requirements across different lifecycle stages, including Invest, Tolerate, Retire, and Migrate.

Stakeholder Management :

- Strong communication and collaboration skills.

- Ability to work effectively with application owners, infrastructure teams, cloud teams, engineering teams, service-management teams, and senior leadership.

Nice-to-Have Skills :

- Familiarity with AI and GenAI architectures.

- Knowledge of Prompt Engineering, Knowledge Graphs, and Retrieval-Augmented Generation (RAG).

- Experience with cloud-native observability frameworks on AWS, Microsoft Azure, and Google Cloud Platform (GCP).

- Exposure to Agentic AI concepts.

- Experience with AI-driven auto-triage or intelligent automation solutions.

- Familiarity with predictive monitoring and AIOps platforms.

Key Success Measures :

- Success in this role will be demonstrated through :

1. A high-performing 14-member Tier 1 AIOps Support team delivering reliable and SLA-compliant triage.

2. Successful expansion of monitoring and triage coverage from 50 to approximately 320 applications.

3. Consistent improvement in Mean-Time-to-Triage (MTTT) and incident-routing accuracy.

4. Reduction in false positives and alert noise.

5. Progressive reduction in manual and reactive triage through automation.

6. Increased adoption of proactive and early-warning monitoring.

7. A standardized and high-quality OpenTelemetry-based telemetry backbone.

8. Strong governance of SLA/SLO, alert, incident, and operational-quality metrics.

9. Comprehensive runbooks and standardized operational processes.

10. A monitoring environment that is ready for future AI-driven and autonomous support capabilities.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...