HamburgerMenu
hirist

Senior Engineer - AI Inference & Kernel Optimization

Square Root Consulting
8 - 15 Years
rupee70-90 LPA
Bangalore

Posted on: 12/08/2026

Job Description

Description :

Position : Sr Engineer AI Inference & Kernel Optimization

Location : Bangalore, India (On-site / Hybrid)

Compensation : ?7090 LPA (Base) + ESOPs + Performance Variable

(Compensation aligned with experience and technical depth)

About the Role :

We are building next-generation AI inference processors optimized for ultra-low latency, high-throughput workloads. As a Senior / Principal Engineer, you will play a critical role in designing and optimizing low-level software and compute kernels that extract maximum performance from GPUs and custom accelerators.

This role is ideal for engineers who thrive close to hardware, enjoy performance tuning, and want to influence the software stack of cutting-edge AI silicon.

Key Responsibilities :

- Design, implement, and optimize high-performance compute kernels for GPUs and/or custom AI accelerators

- Develop low-level software in C/C++ and CUDA, targeting inference workloads for deep learning models

- Apply advanced code optimization techniques, including :

a. Vectorization (SIMD)

b. Memory hierarchy optimization (registers, shared memory, caches)

c. Parallelization strategies

- Cache utilization and memory bandwidth optimization

- Drive profiling, benchmarking, and performance tuning to achieve optimal resource utilization and minimal inference latency

- Collaborate closely with architecture, compiler, and hardware teams to co-design performant solutions

- Analyze bottlenecks across compute, memory, and interconnects, and propose architectural or software improvements

- Mentor junior engineers and contribute to technical direction (for Principal-level candidates)

Required Qualifications :

- Strong expertise in C and C++, with deep understanding of low-level programming

- Hands-on experience with CUDA and GPU programming

- Proven experience developing high-performance kernels

- Deep knowledge of performance optimization techniques, including :

a. Vectorization and instruction-level optimization

b. Threading and parallel execution models

c. Memory hierarchies and cache behavior

- Experience with profiling and performance analysis tools (e.g., Nsight, VTune, perf, custom profilers)

- Strong understanding of AI inference workloads (CNNs, Transformers, GEMM, attention, activation functions, etc.)

Preferred / Nice-to-Have :

- Experience working on AI inference frameworks, runtimes, or compilers

- Background in computer architecture or microarchitecture

- Experience optimizing for latency-critical systems

- Exposure to custom silicon bring-up or hardware-software co-design

- Contributions to performance-critical open-source projects

Level Expectations :

Senior Engineer :

- Own complex kernel implementations

- Independently drive optimization and tuning

- Deliver production-quality, high-performance code

Principal Engineer :

- Define performance strategy and best practices

- Influence architecture and software stack decisions

- Lead complex, cross-functional technical initiatives

Why Join Us :

- Work on cutting-edge AI inference silicon

- Own performance-critical parts of the stack, close to hardware

- High-impact role with strong technical ownership

- Competitive compensation with meaningful ESOP upside


info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...