HamburgerMenu
hirist

Senior AI Infrastructure Engineer

TalentEdge Recruitment Consultants
4 - 8 Years
Vadodara/Baroda

Posted on: 23/07/2026

Job Description

Job Title : Senior AI Infrastructure Engineer (NVIDIA GPU Systems / HPC Environment)

Location : Remote

Experience Required : 5+ years

Time Zone : PST

Position Overview :

We are seeking a highly experienced AI Infrastructure Engineer to support and optimize GPU-based server environments for AI/ML workloads and high-performance computing (HPC). The role involves L3 support, troubleshooting, optimization, and maintenance of enterprise AI compute environments containing high-density GPU servers.

Key Responsibilities :

- Manage and support high-performance AI server infrastructure (192+ CPU cores, 2TB RAM, 8x NVIDIA GPU servers)

- Troubleshoot GPU infrastructure issues (hardware, CUDA, drivers, NVLink, PCIe)

- Support and optimize Mellanox InfiniBand networking for distributed compute operations

- Monitor and maintain AI/HPC clusters for ML training and inference

- Diagnose and resolve Linux system-level issues across GPU compute nodes

- Work with containerized/orchestration technologies (Docker, Kubernetes, Slurm, NVIDIA Container Toolkit)

- Support distributed AI workloads and GPU resource allocation

- Perform system tuning, patching, firmware updates, and performance optimization

- Collaborate with engineering teams to improve scalability and reliability

- Develop documentation, SOPs, and escalation procedures

Required Skills & Experience :

- 5+ years supporting Linux-based enterprise infrastructure

- Strong experience with NVIDIA GPU systems in AI/HPC environments

- Hands-on experience with CUDA and GPU driver management

- Experience supporting GPU clusters or AI training infrastructure

- Knowledge of Mellanox InfiniBand networking

- Strong understanding of distributed compute environments

- Experience with Docker and Kubernetes

- Experience troubleshooting GPU performance bottlenecks

- Familiarity with HPC or AI datacenter infrastructure

Preferred Qualifications :

- Experience with NVIDIA B200/B300, NVLink, RDMA, Slurm, OpenStack

- Experience with AI model training infrastructure

- NVIDIA certifications (preferred)

- RHCE or Kubernetes certifications (a plus)

Soft Skills :

- Strong analytical and troubleshooting abilities

- Ability to work independently in complex technical environments

- Excellent communication and documentation skills

- Comfortable working in fast-paced AI infrastructure environments

Support Level :

- Level 3 (L3) Infrastructure Support

Environment :

- AI Infrastructure / HPC / GPU Datacenter Operations

Please share candidate profiles that match these requirements along with their monthly costing at the earliest.

The job is for:

May work from home
info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...