HamburgerMenu
hirist

GPU Infrastructure Engineer

Skillventory
8 - 15 Years
Multiple Locations

Posted on: 17/06/2026

Job Description

Role Overview :

We are seeking a seasoned GPU Infrastructure Engineer to lead the design, deployment, and optimization of our high-performance computing environments in Hyderabad/Mumbai. In this role, you will serve as the technical backbone for our AI and machine learning initiatives, ensuring our GPU clusters operate at peak efficiency to support large-scale model training and inference. You will collaborate closely with data scientists, software engineers, and IT operations teams to bridge the gap between raw hardware capabilities and high-level application performance. By architecting robust, scalable infrastructure, you will directly influence the speed and reliability of our product delivery, empowering our teams to push the boundaries of computational innovation.


Key Responsibilities :

- Architect and maintain large-scale GPU clusters to provide a stable and high-performance foundation for complex AI workloads.

- Optimize CUDA kernels and system-level configurations to maximize hardware utilization and reduce latency for critical applications.

- Automate infrastructure provisioning and lifecycle management using containerization tools to ensure consistent deployment environments across the organization.

- Implement advanced monitoring and diagnostic frameworks to proactively identify and resolve hardware bottlenecks or performance degradation.

- Partner with cross-functional engineering teams to troubleshoot complex system issues, ensuring seamless integration between hardware resources and software stacks.

- Lead the evaluation and integration of emerging GPU technologies and datacenter hardware to maintain a competitive edge in computational capacity.

Required Skillset :


- Demonstrated expertise in managing high-performance computing (HPC) environments, with a deep understanding of GPU architecture and datacenter operations.

- Proficiency in Python for developing automation scripts and performance monitoring tools that streamline infrastructure management.

- Advanced capability in utilizing Docker and container orchestration platforms to deploy and scale distributed computing workloads efficiently.

- Strong command of CUDA programming and optimization techniques to enhance the performance of compute-intensive applications.

- Exceptional communication skills, with the ability to translate complex technical infrastructure requirements into actionable insights for non-technical stakeholders.

- Proven track record of working in a hybrid or on-site environment, demonstrating the adaptability required to manage physical datacenter assets effectively.

- A Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field, complemented by 8-15 years of hands-on experience in infrastructure engineering.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...