Posted on: 21/07/2026
Role : Support Engineer - AI Factory Network & Connectivity
About the Role :
You will be part of the Tier 2 support team responsible for the day-to-day health of CPT1 AI Factory, a 2,048-GPU NVIDIA H200 cluster built for large-scale AI training and inference. This role is biased toward network and connectivity engineering : keeping the east-west InfiniBand fabric and the north-south Ethernet/WAN paths in peak condition so that GPU training jobs run without interruption. You will work a 247 shift rotation alongside the NOC at Africa Data Centres and escalate to L2/L3 engineering when needed.
Key Responsibilities :
- Monitor, triage, and resolve network-layer incidents across the InfiniBand NDR 400G compute fabric (NVIDIA QM9700 switches) and the Ethernet storage/WAN fabric (NVIDIA Spectrum-4 SN5600, Arista 7280R3 WAN routers).
- Respond to Tier 1 alerts within the defined SLA windows : 4 hours for non-critical, 1 hour for degraded service, and 15 minutes for critical outages.
- Own first-response troubleshooting for connectivity faults : link flaps, NCCL all-reduce degradation, RDMA errors, IB subnet manager issues, and BGP/IP path failures on the LIT MPLS uplinks.
- Perform health checks on the 6 100G WAN links, validate path diversity, and monitor utilisation against the 600G aggregate capacity target.
- Support CSP direct-connect circuits, Azure ExpressRoute, AWS Direct Connect, GCP Partner Interconnect, Oracle Fast Connect provisioned via Teraco.
- Operate and maintain perimeter security : FortiGate 4400 firewall rules, stateful inspection policies, IDPS and TLS inspection where enabled.
- Administer the Rafay orchestration platform for bare-metal network provisioning : VRF/VLAN assignment, bond interfaces, cloud-init templates.
- Maintain and update NOC runbooks; create and close tickets in the incident management system with accurate RCA notes.
- Participate in planned maintenance windows (max 4 hrs/month) and ensure 7-day advance notifications are issued to customers.
- Collaborate with L2/L3 engineers and OEM (HPE/NVIDIA) support on complex IB fabric or WAN routing issues.
Required Skills & Experience :
- Having an engineering degree is a must or MCA or MCS with science in 12th STD.
- 3+ years in a networking or NOC role in a datacentre or HPC/cloud environment.
- Solid understanding of InfiniBand architecture : fat-tree topologies, NDR/HDR speeds, subnet managers (OpenSM/UFM), adaptive routing, RDMA/RoCE, and SHARP in-network compute.
- Hands-on experience with high-speed Ethernet switching : ECMP, VLAN, LAG/MLAG, BGP/OSPF, and MACSEC-encrypted WAN circuits.
- Familiarity with NVIDIA ConnectX NIC configuration (mlxconfig), IB diagnostics (ibstat, ibdiagnet, perfquery), and NCCL test suites.
- Experience with firewalls and network security appliances FortiGate preferred.
- Working knowledge of Linux networking : bonding, VRFs, namespaces, tc/iproute2, and TCP/IP tuning for high-throughput data transfers.
- Comfortable reading Grafana dashboards, Victoria Metrics queries, and interpreting NVIDIA DCGM network counters.
- Excellent incident documentation and communication skills able to write clear P1 updates for both technical and commercial stakeholders.
Nice to Have :
- NVIDIA networking certifications (NCP-MCI or equivalent).
- Experience with GPU cluster fabrics in an AI/ML training environment (NCCL, all-reduce, ring-allreduce).
- Exposure to Slurm job scheduling and how topology-aware scheduling interacts with IB fabric health.
- Scripting skills in Python or Bash for automation of network health checks.
- Familiarity with Ansible/NautoBot for network automation and CMDB management.
Did you find something suspicious?
Posted by
Liquid Intelligent Technologies
HR at Liquid Intelligent Technologies
Last Active: NA as recruiter has posted this job through third party tool.
Posted in
AI/ML
Functional Area
Networking & Wireless
Job Code
1656264