Posted on: 02/07/2026
Roles & Responsibilities :
Responsibilities :
- Monitor production infrastructure, platform health, and application uptime across 24x7 environments, ensuring SLA adherence and rapid incident response.
- Detect, triage, and escalate incidents in real time, coordinating with engineering and DevOps teams to drive resolution within defined RTO and RCA timelines.
- Manage and respond to alerts from monitoring tools (Grafana, PagerDuty, Datadog, or equivalent), distinguishing signal from noise and reducing MTTR.
- Execute routine operational tasks including deployments, configuration changes, log analysis, and scheduled maintenance activities.
- Maintain and improve runbooks, escalation playbooks, and incident documentation to build institutional knowledge and reduce repeat issues.
- Collaborate with the engineering team on observability improvements, including alert tuning, dashboard creation, and proactive capacity monitoring.
- Participate in post-incident reviews and contribute to root cause analysis and preventive action planning.
Ideal candidate :
- 2-3 years of experience in a NOC, infrastructure operations, or DevOps support role.
- Solid foundation in Computer Science, with hands-on understanding of networking, Linux systems, and cloud infrastructure.
- Proficiency with monitoring and observability tools such as Grafana, Prometheus, Datadog, Graylog, or similar.
- Working knowledge of containerisation and orchestration technologies, Docker and Kubernetes preferred.
- Comfort with scripting in Bash, Python, or similar for automation and operational tasks.
- Familiarity with CI/CD pipelines and DevOps workflows is a strong plus.
- Strong communication skills, ability to write clear incident updates and escalation notes under pressure.
Did you find something suspicious?