Posted on: 11/09/2026
Role Overview :
- 3+ years of experience in application production support with a strong track record of independently diagnosing and resolving incidents.
Technical Requirements :
- Solid working knowledge of the full technology stack in scope including event streaming platforms, integration middleware, AKS-hosted microservices, and observability tooling.
- Hands-on experience with incident lifecycle management in ticketing systems (iTrack or equivalent), including root cause identification and resolution documentation.
- Proficient with Splunk, PagerDuty, Prometheus, and Grafana for active troubleshooting and issue resolution.
- Hands-on operational experience with Kubernetes, especially Azure Kubernetes Service (AKS), including pod-level diagnostics, restarts, and health investigation.
- Practical working knowledge of Confluent Kafka and Azure Event Hub : consumer lag analysis, topic health checks, and message flow troubleshooting.
- Solid SQL/Postgres skills for data-level investigation and validation during incidents.
- Working ability to read and interpret Java, Spring Boot, and React application logs for issue identification.
- Basic Python scripting capability for operational checks and quick-fix automation.
- Good Linux/Unix command-line skills for real-time log analysis and system diagnostics.
Soft Skills & Shift Requirements :
- Strong written and verbal communication skills for incident updates, resolution documentation, and client coordination.
- Willingness to work in rotational 24x7 shifts.
Requirement :
- Only Immediate Joiners are considerable.
Did you find something suspicious?
Posted by
Posted in
DevOps / SRE
Functional Area
Systems Administration
Job Code
1670630