HamburgerMenu
hirist

GPU Kernel Developer/AI Performance Engineer

SOFTPATH TECHNOLOGIES
4 - 8 Years
Multiple Locations

Posted on: 08/05/2026

Job Description

Job Title : GPU Kernel Developer / AI Performance Engineer


Company: Wipro

Employment Type : Full-Time (FTE)


Experience : 4- 8 Years


Location : Hyderabad / Bangalore / Chennai / Pune / Remote


Job Summary :


We are looking for highly skilled GPU Kernel Developers / AI Performance Engineers with strong expertise in Python, C/C++, PyTorch, Triton, and GPU kernel development. The ideal candidate should have hands-on experience in optimizing AI/ML workloads, developing high-performance GPU kernels, and porting machine

learning models across different hardware platforms.

This role requires deep knowledge of GPU architecture, performance optimization, compiler/runtime execution layers, and low-level systems programming to improve AI model efficiency and scalability.


Key Responsibilities :


- Design, develop, and optimize GPU kernels for high-performance AI/ML workloads.

- Develop custom kernels using Triton for accelerating deep learning operations.

- Work on GEMM kernel development (matrix multiplication optimization).

- Port machine learning models to new hardware platforms and ensure compatibility.

- Optimize GPU compute performance by improving memory utilization, latency, and throughput.

- Perform system-level performance tuning, benchmarking, and stress testing.

- Work with compiler/runtime execution layers for model optimization.

- Develop and integrate custom SDKs or hardware abstraction layers.

- Collaborate with AI/ML teams to improve model inference and training performance.

- Troubleshoot low-level performance bottlenecks in GPU workloads.


Required Skills :


Programming :


- Strong expertise in Python


- Strong hands-on experience in C/C++

- Experience with PyTorch

- Hands-on expertise in Triton language/kernel development

GPU & System-Level Expertise :


- Mandatory experience in GPU kernel development

- Strong understanding of GPU architecture

- Experience in compute optimization

- Experience with compiler optimizations/runtime execution layers

- Experience working with custom SDKs/HAL layers

Performance Optimization :


- Hands-on experience in GEMM kernel development

- Experience in ML model porting

- Strong experience in system-level performance tuning

- Stress testing and benchmarking expertise


Preferred Skills :


- Exposure to CUDA/OpenCL/ROCm is a plus

- Experience in AI acceleration frameworks

- Knowledge of deep learning model optimization techniques


- Experience with distributed systems or large-scale AI workloads

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...