Posted on: 11/09/2026
About the Role :
We are seeking an experienced and dynamic AI Support Lead / Manager to oversee a modern, standardized observability and monitoring capability across a large enterprise application landscape.
In this role, you will drive operational excellence across a team of Tier-1 Support Engineers, taking initial responsibility for triage support across ~50 applications and scaling to ~320 applications over the next two years.
You will own the operational discipline, process hygiene, and telemetry data quality necessary to drive reliable triage while transitioning operations from reactive manual triage toward automated, proactive monitoring as our OpenTelemetry backbone matures.
Key Responsibilities :
1. Operations & Triage Quality :
- Set and enforce strict standards for alert detection, classification, and routing to optimize Mean-Time-to-Triage (MTTT) and ensure adherence to SLAs/SLOs.
2. Application Onboarding :
- Execute a repeatable, structured framework to seamlessly scale application coverage from 50 to 320+ enterprise applications.
3. Incident & Change Management :
- Direct outage response workflows, lead cross-functional coordination during major incidents, and drive robust Root Cause Analysis (RCA) and runbook creation.
4. Shift & Roster Governance :
- Coordinate weekly 24x7 rotational shift coverage, ensuring continuous operational availability across global time zones.
5. Telemetry & Automation Readiness :
- Partner with application teams to standardize heterogeneous logs into an OpenTelemetry backbone.
- Drive continuous reduction of false positives and manual overhead to prepare the ecosystem for Agentic AI auto-triage tools.
6. Executive Reporting :
- Track, analyze, and present SLA/SLO metrics, alert accuracy, and operational health reports to senior program leadership.
Must-Have Requirements :
- Experience : 8 - 12 years in IT Application Support, Incident Management, or AIOps environment with strong lead/techno-functional ownership experience.
- ITIL & Governance : Deep mastery of Incident Management lifecycles, Problem/Change Management, CMDB concepts, and SLA/SLO calculation frameworks.
- Observability Tooling : Hands-on experience with modern monitoring platforms (e.g., Dynatrace, New Relic, AWS CloudWatch, ManageEngine, Glassbox) and exposure to OpenTelemetry or distributed tracing.
- Infrastructure & Cloud : Core understanding of VMs, load balancers, firewalls, cloud infrastructure (AWS/Azure/GCP), containers, OpenShift (OCP), and Kubernetes.
- Scripting & APIs : Proficiency with Shell scripting/batch files, JSON/XML data formats, and API testing tools (Postman, SOAP UI).
- Shift Flexibility : Readiness to work in a 24x7 rotational shift setup (including IST and night coverage).
Good-to-Have Qualifications :
- Familiarity with modern AI concepts, including Prompt Engineering, RAG architectures, Knowledge Graphs, or Agentic AI automation tools.
- Experience managing cloud-native observability frameworks in large-scale enterprise environments.
- Prior experience in product/software development setups navigating complex application lifecycles.
Interview Process :
- Round 1 : Virtual Technical Screening
- Round 2 : Face-to-Face Technical & Operational Discussion
Why Join Us? :
- High-Impact Domain : Direct oversight of enterprise-scale observability and AI-driven support transformation.
- Cutting-Edge Tech : Hands-on involvement in implementing OpenTelemetry and next-gen AI automation workflows.
- Flexible Engagement : Supportive hybrid work model with rotational shift flexibility and comprehensive employee benefits
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
IT Management / IT Support
Job Code
1670705