HamburgerMenu
hirist

Gnani.ai - LLMOps Engineer

Gnani Innovations
6 - 10 Years
Bangalore

Posted on: 08/10/2026

Job Description

Role Overview:

We are hiring an LLMOps Engineer to own how our large language models actually run in production - the serving stack, the optimization work that makes it affordable, and the low-level performance engineering that makes it fast. This is a deep infrastructure and performance role, not a wrapper-and-API role.

You will work on self-hosted foundational models served on large-scale NVIDIA GPU clusters, under real-time latency budgets set by live spoken conversations. Voice inference is unforgiving: time-to-first-token is a product feature, not a number on a dashboard, and every millisecond of tail latency is audible to a caller. Your mandate is to hold latency and quality while driving cost per million tokens down.

The work spans three layers: the serving engine (vLLM, SGLang, NVIDIA Dynamo), the model itself (quantization, distillation, speculative decoding, pruning), and the kernels underneath (CUDA/Triton, attention and MoE kernels, profiling and bottleneck analysis). We expect real depth in at least two of the three and genuine curiosity about the third.

Key Responsibilities:

Inference Serving and Platform:

- Own the production LLM serving stack across vLLM, SGLang, and NVIDIA Dynamo - including engine selection per workload, with benchmark evidence to justify it

- Tune the serving path end to end: continuous batching, chunked prefill, paged and radix attention, prefix and KV caching, cache-aware and sticky routing, speculative decode integration

- Design and operate disaggregated prefill/decode and multi-node deployments; configure tensor, pipeline, and expert parallelism for both dense and Mixture-of-Experts models

- Run inference on Kubernetes with autoscaling, rolling model updates, canary releases, and tested rollback paths; enforce staging-to-production promotion gates

- Instrument everything that matters: TTFT, inter-token latency, p50/p95/p99, tokens/sec/GPU, KV cache hit rate, queue depth, GPU utilization and MFU, and cost per million tokens

- Build and maintain internal serving APIs, model registries, and deployment tooling so that model updates are routine rather than events

Model Optimization:

- Quantization: FP8, INT8, and INT4 weight and activation quantization plus KV cache quantization - calibration set design, per-layer sensitivity analysis, and accuracy recovery, with measured deltas per language and task

- Speculative decoding: draft-model, EAGLE/Medusa-style, and n-gram approaches; tuned for real acceptance rate and end-to-end latency gain under production traffic, not theoretical speedup

- Distillation: build smaller task-specific students from larger teachers for latency-critical paths, held to explicit eval-parity targets

- Pruning and adapters: structured pruning, LoRA/adapter serving, and multi-adapter batching for per-deployment specialization

- Compilation: torch.compile, CUDA graphs, and TensorRT-LLM engine builds - including diagnosing recompilation and dynamic-shape stalls

Kernel and Low-Level Performance:

- Profile with Nsight Systems/Compute and the PyTorch profiler; classify bottlenecks as memory-bound, compute-bound, or launch/synchronization-bound, and act on the classification

- Write and tune custom kernels in CUDA and Triton - attention variants, fused MoE dispatch, sampling, quantized GEMM

- Eliminate host-device synchronization stalls, dynamic-shape recompilation, and unoverlapped collectives in distributed serving

- Work fluently with attention and MoE kernel libraries (FlashAttention, FlashInfer, CUTLASS/cuBLAS, expert-parallel dispatch libraries) and know when to use them versus write your own

- Reason from first principles about arithmetic intensity, memory bandwidth, and roofline limits before reaching for a tool

Evaluation, Reliability, and Cost:

- Own the optimization regression gate: no optimized build reaches production without an accuracy and behavior evaluation across languages and task types

- Build load-testing harnesses that replay realistic traffic - concurrency, length distributions, burstiness, multi-turn sessions

- Run capacity planning and cost modeling; set and hit targets for cost per million tokens and per concurrent session

- Carry on-call for inference services, write runbooks, and lead blameless postmortems on latency and availability incidents

- Work closely with the training and post-training teams so that serving constraints inform model architecture decisions early, not after the checkpoint lands

Must Have:

- 6 - 10 years total experience, with 3+ years owning large-scale LLM or speech inference in production

- Has owned an inference platform end to end: architecture, SLOs, capacity, cost, deployment safety, and on-call

- Has written or substantially tuned custom CUDA or Triton kernels that shipped to production

- MoE serving at scale: expert parallelism, routing load imbalance, dispatch kernels, and the failure modes specific to sparse models

- Deep profiling ability - has diagnosed and fixed a utilization or MFU collapse in a distributed multi-node system, and can walk through the investigation

- Track record running an accuracy-preserving optimization program (quantization plus speculative decoding plus distillation) behind real evaluation gates

- Sets technical direction, mentors engineers, and can make a defensible build-versus-adopt call on serving infrastructure

- Upstream contributions to open-source serving or kernel projects (vLLM, SGLang, FlashInfer, TensorRT-LLM, or similar) - a strong plus, and close to an expectation at this level

Good to Have:

- Real-time speech inference: streaming ASR/TTS, sub-second time-to-first-audio budgets, barge-in and turn-taking constraints

- Serving hybrid-architecture models (state-space/Mamba blocks combined with attention) and understanding their distinct state and cache management

- TensorRT-LLM engine building, NVIDIA NIM, or Triton Inference Server in production

- Experience with Indic or other multilingual and code-mixed models, and the evaluation discipline that comes with them

- Ray Serve, KServe, LLM gateways, or semantic and prefix cache layers

- Enough training-side exposure (Megatron, NeMo, DeepSpeed, FSDP) to work fluently with the pretraining team

- Public benchmarks, blog posts, or talks on inference optimization

What You Will Work On:

You will work on in-house foundational models - not third-party API endpoints - running on large NVIDIA GPU clusters and serving live enterprise and government-scale voice traffic. The models are ours, so the whole stack is open to you: you can change the serving engine, the quantization recipe, the kernel, or the model itself, and you will often need to change more than one to hit a target.

The constraints are real and the feedback loop is fast. A latency regression is heard by callers within minutes. A successful optimization shows up directly in infrastructure spend. You will have access to proprietary speech and language data, in-house training and evaluation infrastructure, a close engineering partnership with NVIDIA, and a team that has been building Indian-language AI since 2016.

This role has a clear path to owning the inference platform and its technical direction, or to a deeper specialization in performance and kernel engineering, depending on where your strength lies.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...