Posted on: 25/09/2026
We are seeking an experienced and dynamic AIOps Support Lead (Manager) to lead a modern, standardized observability and monitoring capability across a diverse enterprise application landscape.
The role will manage a 14-member Tier 1 AIOps Support team, initially supporting approximately 50 applications and scaling to nearly 320 applications over the next two years. The position will own people management, operational processes, triage quality, telemetry/data discipline, and continuous automation initiatives.
The key objective is to transition the organization from reactive, manual incident triage toward proactive and increasingly automated monitoring, supported by a mature OpenTelemetry-based observability framework.
Key Responsibilities :
1. Team & Shift Management :
- Lead, hire, coach, develop, and manage a team of 14 Tier 1 AIOps Support Engineers.
- Manage team performance, productivity, capability development, and resource allocation.
- Design and maintain effective 24x7 shift and roster coverage.
- Scale team operations in line with application portfolio growth.
- Establish clear roles, responsibilities, operating standards, and performance expectations.
2. Operations & Triage Management :
- Own the quality, consistency, and speed of incident detection, classification, and routing.
- Establish and monitor standards for Mean-Time-to-Triage (MTTT).
- Develop a structured and repeatable framework for onboarding applications into the AIOps monitoring ecosystem.
- Scale application onboarding from approximately 50 to 320 applications.
- Lead major outage and incident response activities.
- Coordinate with application, infrastructure, cloud, and other technical teams during critical incidents.
- Drive Root Cause Analysis (RCA) following major incidents.
- Develop and continuously improve technical runbooks, operating procedures, and knowledge repositories.
3. Observability & Data Quality :
- Establish standards for ingesting and managing heterogeneous telemetry and log data.
- Partner with application and technology teams to improve the quality, consistency, and usability of monitoring data.
- Support the transition toward a standardized OpenTelemetry backbone.
- Improve alert classification, correlation, enrichment, and prioritization.
- Reduce alert noise and false positives through continuous monitoring and optimization.
- Establish strong data-quality governance to ensure reliable automated triage.
4. Automation & Proactive Monitoring :
- Drive continuous reduction of manual triage activities through automation.
- Identify opportunities for automated alert correlation, event enrichment, incident routing, and early-warning detection.
- Support the transition from reactive monitoring to proactive and predictive operations.
- Prepare the observability environment for future AI-driven and Agentic AI automation.
- Collaborate with engineering and platform teams to identify automation use cases and improve operational efficiency.
5. Incident, Problem & Change Management :
- Ensure adherence to ITIL-based Incident, Problem, and Change Management processes.
- Drive effective incident escalation and resolution workflows.
- Ensure appropriate linkage between incidents, problems, changes, configuration items, and business services.
- Maintain alignment with CMDB and service-management practices.
- Monitor adherence to SLA/SLO commitments and operational governance standards.
6. KPI & Governance Reporting :
- Define and track operational KPIs and service-quality metrics.
- Prepare regular governance reports for program and technology leadership.
- Monitor metrics including :
1. SLA/SLO compliance
2. Mean-Time-to-Triage (MTTT)
3. Alert-to-incident accuracy
4. False-positive rate
5. Incident volume and trends
6. Automation percentage
7. Manual triage reduction
8. Application onboarding progress
- Present insights, trends, risks, and improvement opportunities to senior stakeholders.
Technical Requirements :
IT Service Management & Governance :
- Strong understanding of ITIL Incident Management lifecycle.
- Knowledge of Problem Management and Change Management.
- Understanding of CMDB concepts and service relationships.
- Strong understanding of SLA/SLO definitions and availability calculations.
- Experience working within structured operational governance frameworks.
Observability & Monitoring :
- Hands-on experience with one or more modern monitoring and observability platforms, such as Dynatrace, New Relic, AWS CloudWatch, ManageEngine, Glassbox, or similar enterprise monitoring/observability platforms.
- Additional exposure to OpenTelemetry, distributed tracing, Application Performance Monitoring (APM), log aggregation, metrics and event monitoring, and alert correlation and enrichment.
Infrastructure & Networking :
- Good understanding of Virtual Machines (VMs), firewalls, load balancers, containers, Kubernetes, OpenShift / OCP, and enterprise infrastructure and application architecture.
Scripting, APIs & Data Formats :
- Working knowledge of Unix/Linux shell scripting and exposure to Windows batch scripting.
- Understanding of JSON and XML data formats.
- Experience with API testing and troubleshooting tools such as Postman, SOAP UI, or similar API testing tools.
Cloud & Security :
- Understanding of cloud fundamentals including compute, storage, networking, and cloud monitoring.
- Basic understanding of security technologies and protocols including TLS, SSL, authentication tokens, and secrets management.
- Awareness of data infrastructure concepts such as data lineage, data latency, and data quality.
Leadership & Behavioral Competencies :
Team Leadership :
- Proven experience managing and scaling technical support or operations teams.
- Strong people-management, coaching, mentoring, and performance-management skills.
- Ability to manage teams operating in 24x7 support environments.
Scale-Up & Transformation :
- Demonstrated experience scaling support operations during periods of significant application or business growth.
- Ability to establish standardized processes while maintaining operational flexibility.
Analytical & Problem-Solving Skills :
- Ability to analyze fragmented and high-volume alert data.
- Strong capability to convert noisy monitoring signals into actionable insights.
- Ability to develop effective technical runbooks and operational procedures.
Adaptability :
- Comfortable operating in complex and ambiguous environments.
- Ability to dynamically prioritize applications and support requirements across different lifecycle stages, including Invest, Tolerate, Retire, and Migrate.
Stakeholder Management :
- Strong communication and collaboration skills.
- Ability to work effectively with application owners, infrastructure teams, cloud teams, engineering teams, service-management teams, and senior leadership.
Nice-to-Have Skills :
- Familiarity with AI and GenAI architectures.
- Knowledge of Prompt Engineering, Knowledge Graphs, and Retrieval-Augmented Generation (RAG).
- Experience with cloud-native observability frameworks on AWS, Microsoft Azure, and Google Cloud Platform (GCP).
- Exposure to Agentic AI concepts.
- Experience with AI-driven auto-triage or intelligent automation solutions.
- Familiarity with predictive monitoring and AIOps platforms.
Key Success Measures :
- Success in this role will be demonstrated through :
1. A high-performing 14-member Tier 1 AIOps Support team delivering reliable and SLA-compliant triage.
2. Successful expansion of monitoring and triage coverage from 50 to approximately 320 applications.
3. Consistent improvement in Mean-Time-to-Triage (MTTT) and incident-routing accuracy.
4. Reduction in false positives and alert noise.
5. Progressive reduction in manual and reactive triage through automation.
6. Increased adoption of proactive and early-warning monitoring.
7. A standardized and high-quality OpenTelemetry-based telemetry backbone.
8. Strong governance of SLA/SLO, alert, incident, and operational-quality metrics.
9. Comprehensive runbooks and standardized operational processes.
10. A monitoring environment that is ready for future AI-driven and autonomous support capabilities.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
IT Management / IT Support
Job Code
1674860