Posted on: 23/07/2026
Job Title : Senior AI Infrastructure Engineer (NVIDIA GPU Systems / HPC Environment)
Location : Remote
Experience Required : 5+ years
Time Zone : PST
Position Overview :
We are seeking a highly experienced AI Infrastructure Engineer to support and optimize GPU-based server environments for AI/ML workloads and high-performance computing (HPC). The role involves L3 support, troubleshooting, optimization, and maintenance of enterprise AI compute environments containing high-density GPU servers.
Key Responsibilities :
- Manage and support high-performance AI server infrastructure (192+ CPU cores, 2TB RAM, 8x NVIDIA GPU servers)
- Troubleshoot GPU infrastructure issues (hardware, CUDA, drivers, NVLink, PCIe)
- Support and optimize Mellanox InfiniBand networking for distributed compute operations
- Monitor and maintain AI/HPC clusters for ML training and inference
- Diagnose and resolve Linux system-level issues across GPU compute nodes
- Work with containerized/orchestration technologies (Docker, Kubernetes, Slurm, NVIDIA Container Toolkit)
- Support distributed AI workloads and GPU resource allocation
- Perform system tuning, patching, firmware updates, and performance optimization
- Collaborate with engineering teams to improve scalability and reliability
- Develop documentation, SOPs, and escalation procedures
Required Skills & Experience :
- 5+ years supporting Linux-based enterprise infrastructure
- Strong experience with NVIDIA GPU systems in AI/HPC environments
- Hands-on experience with CUDA and GPU driver management
- Experience supporting GPU clusters or AI training infrastructure
- Knowledge of Mellanox InfiniBand networking
- Strong understanding of distributed compute environments
- Experience with Docker and Kubernetes
- Experience troubleshooting GPU performance bottlenecks
- Familiarity with HPC or AI datacenter infrastructure
Preferred Qualifications :
- Experience with NVIDIA B200/B300, NVLink, RDMA, Slurm, OpenStack
- Experience with AI model training infrastructure
- NVIDIA certifications (preferred)
- RHCE or Kubernetes certifications (a plus)
Soft Skills :
- Strong analytical and troubleshooting abilities
- Ability to work independently in complex technical environments
- Excellent communication and documentation skills
- Comfortable working in fast-paced AI infrastructure environments
Support Level :
- Level 3 (L3) Infrastructure Support
Environment :
- AI Infrastructure / HPC / GPU Datacenter Operations
Please share candidate profiles that match these requirements along with their monthly costing at the earliest.
The job is for:
Did you find something suspicious?