Posted on: 25/08/2026
Job Description :
We are looking for a Senior FDE with strong expertise in AI infrastructure, GPU clusters, Kubernetes, and distributed systems to work directly with customers on production AI workloads.
Key Responsibilities :
- Design and optimize GPU infrastructure and LLM inference platforms.
- Work with NVIDIA/AMD GPUs, CUDA/ROCm, vLLM, TensorRT-LLM, and SGLang.
- Deploy and manage Kubernetes-based AI infrastructure.
- Automate infrastructure using Terraform, Ansible, and Helm.
- Troubleshoot GPU drivers, NCCL/RCCL, networking, and storage issues.
- Collaborate directly with customer engineering teams to deliver production AI solutions.
Required Skills :
- 6+ years of experience in AI Infrastructure, FDE, Distributed Systems, or Technical Consulting.
- Strong Linux, Kubernetes, Python, and Go skills.
- Hands-on experience with NVIDIA/AMD GPU infrastructure and CUDA/ROCm.
- Knowledge of LLM inference, GPU optimization, RDMA/InfiniBand/RoCE.
- Experience with Terraform/Helm and production cloud infrastructure.
- Strong customer-facing and problem-solving skills.
- Preferred : Experience with production AI systems, GPU vendors/cloud platforms, and high-performance computing environments.
Role Details :
- Role : Back End Developer
- Industry Type : Software Product
- Department : Engineering - Software & QA
- Employment Type : Full Time, Permanent
- Role Category : Software Development
Key Skills : Golang, FDE, GPU, AI Infrastructure, Python, RDMA, ROCm, vLLM, NVIDIA, NVSwitch, Cuda, Infiniband, GPU Cluster, NVLink, NCCL, Kubernetes.
About Company :
HiringEye started by Ex Amazonian which is located in Hyderabad. At HiringEye, we are dedicated to finding exceptional leaders who drive organizational success. We are helping more than 35 plus startups and large scale companies to build their tech team.
This requirement is for one of our client.
Did you find something suspicious?