Posted on: 18/09/2026
NOC Server Engineer
The NOC Server Engineer provides advanced 24x7 operational support for cloud-hosted server infrastructure in AWS and Azure. This role acts as the shift escalation lead for complex and high-severity incidents, ensuring service stability and rapid restoration. Kubernetes (EKS), Terraform, and foundational cloud networking knowledge are required skill sets to support modern cloud workloads.
Key Responsibilities:
- Provide deep troubleshooting for Linux and Windows servers hosted in AWS and Azure, including OS, services, performance, and capacity.
- Analyze alerts and telemetry using CloudWatch, Azure Monitor, Splunk, and Grafana to validate root cause and recovery.
- Coordinate with Operations Center (OC), Incident Managers, and platform SMEs; maintain clear ownership and escalation.
- Execute safe, pre-approved operational actions and validate post-change health using documented procedures.
- Ensure accurate incident timelines, updates, and handovers in ServiceNow/Jira.
- Mentor engineers and own structured shift handovers to maintain operational continuity.
Required Skill Set:
- Cloud Server Engineering: Strong hands-on experience supporting AWS and/or Azure compute services (EC2, Azure VMs/VMSS).
- Linux & Windows Servers: Advanced OS-level administration, patching, log analysis, service troubleshooting, and performance tuning.
- Kubernetes (EKS) & Container Orchestration: Experience triaging EKS node and pod issues, crash loops, scaling symptoms, and workload health.
- Infrastructure-as-Code (Terraform): Ability to execute, review, and validate Terraform-based changes following operational guardrails.
- Cloud Networking Fundamentals: Working knowledge of VPC/VNet, subnets, routing tables, security groups/NSGs, load balancers, DNS, and basic connectivity troubleshooting.
- Observability: CloudWatch, Azure Monitor, Splunk, Grafana/Prometheus; alert tuning and dashboard interpretation.
- ITSM & Operations: Strong incident documentation, communication, and SLA discipline using ServiceNow or Jira.
Required Experience:
- 6+ years of experience in Cloud Operations, NOC, SRE, or Infrastructure Support roles.
- Proven experience leading or handling P1/P2 incidents in 24x7 production environments.
- Hands-on experience working with runbooks, escalation matrices, and shift handovers.
Work Arrangement:
Work from Office with structured handovers and escalation ownership.
Did you find something suspicious?