HamburgerMenu
hirist

AI Operations Support Engineer

Magnet HR Consulting Services
2 - 5 Years
Bangalore

Posted on: 04/09/2026

Job Description

Job Description:

We are seeking an experienced and detail-oriented AIOps Support Engineer to join our growing global technology team. In this role, you will perform hands-on, manual triage of alarms and monitoring alerts across a large and diverse application portfolio.

The portfolio starts with 50 applications and scales to approximately 320 over two years, spanning various technology stacks, log aggregation tools, and lifecycle dispositions (Invest, Tolerate, Retire, Migrate). The immediate focus is disciplined, high-quality manual triage with a clear roadmap toward automated, proactive detection as our underlying telemetry backbone (built on OpenTelemetry) matures. This is a foundational human layer that makes accurate, standardized data and future automation possible.

Key Responsibilities:

- Manual Tier 1 Triage: Monitor, detect, classify, and route incidents based on alarm and alert signals from a range of source applications.

- Multi-Tool Observability: Interpret alerts from multiple log aggregation platforms (Dynatrace, New Relic, ManageEngine, Glassbox, etc.) and translate them into consistent, actionable incident data.

- Telemetry and Data Stewardship: Help build and maintain a clean OpenTelemetry-based data backbone; flag inconsistent, noisy, or low-quality alert data to drive it toward a standardized state.

- Proactive Monitoring: Contribute to the shift from reactive incident response toward proactive issue prevention by flagging patterns and gaps in monitoring coverage.

- Ecosystem Tracking: Learn and track the technology stack, architecture disposition, and ownership of each supported application within a growing portfolio.

- Cross-functional Coordination: Coordinate directly with the correct support groups and application owners to resolve or escalate incidents quickly.

Key Technical Competencies:

- Application and Production Support (L1/L1.5)

- Incident, Problem and Change Management (ITIL)

- Application, Middleware and Database Log Analysis

- Monitoring and Observability Tools (Dynatrace, New Relic, AWS CloudWatch, etc.)

- OpenTelemetry and Distributed Tracing (Working knowledge)

- Linux and Windows OS Troubleshooting

- Basic Networking (TCP/IP, DNS, HTTP/HTTPS, SSL, Load Balancers)

- SQL and Database Query Analysis

- API and Integration Troubleshooting (Postman, SOAP UI, etc.)

- Hybrid Environment Support (On-Premises + Cloud)

- ServiceNow / Jira Ticket Management

Must-Have Skills:

- 2 - 5 years of experience in application or production support (L1/L1.5), preferably in hybrid environments.

- Strong understanding of the Incident Management Lifecycle, and ITIL concepts (Problem, Change Request, Service Request).

- Basic CMDB concepts.

- Hands-on experience with log aggregation technologies.

- Working knowledge of JSON and XML, along with basic file/task automation.

- Knowledge of IT infrastructure: VMs, firewalls, load balancers, containers, OpenShift (OCP), Kubernetes.

- Unix Shell scripting and Windows batch file creation.

- Basic understanding of TLS, SSL, tokens, and secret management.

- Ability to support outage response, contribute to Root Cause Analyses (RCAs), and create runbooks.

- Basic understanding of SLA/SLO concepts and data concepts (data latency, lineage, etc.).

Nice-to-Have Skills:

- Familiarity with AI concepts such as prompt engineering, knowledge graphs, and Retrieval-Augmented Generation (RAG).

- Exposure to public cloud platforms (AWS, Azure, or GCP).

- Scripting or automation experience beyond basic file automation (e.g., Python).

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...