Role Overview:
We are hiring an LLMOps Engineer to own how our large language models actually run in production - the serving stack, the optimization work that makes it affordable, and the low-level performance engineering that makes it fast. This is a deep infrastructure and performance role, not a wrapper-and-API role.
You will work on self-hosted foundational models served on large-scale NVIDIA GPU clusters, under real-time latency budgets set by live spoken conversations. Voice inference is unforgiving: time-to-first-token is a product feature, not a number on a dashboard, and every millisecond of tail latency is audible to a caller. Your mandate is to hold latency and quality while driving cost per million tokens down.
The work spans three layers: the serving engine (vLLM, SGLang, NVIDIA Dynamo), the model itself (quantization, distillation, speculative decoding, pruning), and the kernels underneath (CUDA/Triton, attention and MoE kernels, profiling and bottleneck analysis). We expect real depth in at least two of the three and genuine curiosity about the third.
Key Responsibilities:
Inference Serving and Platform:
- Own the production LLM serving stack across vLLM, SGLang, and NVIDIA Dynamo - including engine selection per workload, with benchmark evidence to justify it
- Tune the serving path end to end: continuous batching, chunked prefill, paged and radix attention, prefix and KV caching, cache-aware and sticky routing, speculative decode integration
- Design and operate disaggregated prefill/decode and multi-node deployments; configure tensor, pipeline, and expert parallelism for both dense and Mixture-of-Experts models
- Run inference on Kubernetes with autoscaling, rolling model updates, canary releases, and tested rollback paths; enforce staging-to-production promotion gates
- Instrument everything that matters: TTFT, inter-token latency, p50/p95/p99, tokens/sec/GPU, KV cache hit rate, queue depth, GPU utilization and MFU, and cost per million tokens
- Build and maintain internal serving APIs, model registries, and deployment tooling so that model updates are routine rather than events
Model Optimization:
- Quantization: FP8, INT8, and INT4 weight and activation quantization plus KV cache quantization - calibration set design, per-layer sensitivity analysis, and accuracy recovery, with measured deltas per language and task
- Speculative decoding: draft-model, EAGLE/Medusa-style, and n-gram approaches; tuned for real acceptance rate and end-to-end latency gain under production traffic, not theoretical speedup
- Distillation: build smaller task-specific students from larger teachers for latency-critical paths, held to explicit eval-parity targets
- Pruning and adapters: structured pruning, LoRA/adapter serving, and multi-adapter batching for per-deployment specialization
- Compilation: torch.compile, CUDA graphs, and TensorRT-LLM engine builds - including diagnosing recompilation and dynamic-shape stalls
Kernel and Low-Level Performance:
- Profile with Nsight Systems/Compute and the PyTorch profiler; classify bottlenecks as memory-bound, compute-bound, or launch/synchronization-bound, and act on the classification
- Write and tune custom kernels in CUDA and Triton - attention variants, fused MoE dispatch, sampling, quantized GEMM
- Eliminate host-device synchronization stalls, dynamic-shape recompilation, and unoverlapped collectives in distributed serving
- Work fluently with attention and MoE kernel libraries (FlashAttention, FlashInfer, CUTLASS/cuBLAS, expert-parallel dispatch libraries) and know when to use them versus write your own
- Reason from first principles about arithmetic intensity, memory bandwidth, and roofline limits before reaching for a tool
Evaluation, Reliability, and Cost:
- Own the optimization regression gate: no optimized build reaches production without an accuracy and behavior evaluation across languages and task types
- Build load-testing harnesses that replay realistic traffic - concurrency, length distributions, burstiness, multi-turn sessions
- Run capacity planning and cost modeling; set and hit targets for cost per million tokens and per concurrent session
- Carry on-call for inference services, write runbooks, and lead blameless postmortems on latency and availability incidents
- Work closely with the training and post-training teams so that serving constraints inform model architecture decisions early, not after the checkpoint lands
Must Have:
- 6 - 10 years total experience, with 3+ years owning large-scale LLM or speech inference in production
- Has owned an inference platform end to end: architecture, SLOs, capacity, cost, deployment safety, and on-call
- Has written or substantially tuned custom CUDA or Triton kernels that shipped to production
- MoE serving at scale: expert parallelism, routing load imbalance, dispatch kernels, and the failure modes specific to sparse models
- Deep profiling ability - has diagnosed and fixed a utilization or MFU collapse in a distributed multi-node system, and can walk through the investigation
- Track record running an accuracy-preserving optimization program (quantization plus speculative decoding plus distillation) behind real evaluation gates
- Sets technical direction, mentors engineers, and can make a defensible build-versus-adopt call on serving infrastructure
- Upstream contributions to open-source serving or kernel projects (vLLM, SGLang, FlashInfer, TensorRT-LLM, or similar) - a strong plus, and close to an expectation at this level
Good to Have:
- Real-time speech inference: streaming ASR/TTS, sub-second time-to-first-audio budgets, barge-in and turn-taking constraints
- Serving hybrid-architecture models (state-space/Mamba blocks combined with attention) and understanding their distinct state and cache management
- TensorRT-LLM engine building, NVIDIA NIM, or Triton Inference Server in production
- Experience with Indic or other multilingual and code-mixed models, and the evaluation discipline that comes with them
- Ray Serve, KServe, LLM gateways, or semantic and prefix cache layers
- Enough training-side exposure (Megatron, NeMo, DeepSpeed, FSDP) to work fluently with the pretraining team
- Public benchmarks, blog posts, or talks on inference optimization
What You Will Work On:
You will work on in-house foundational models - not third-party API endpoints - running on large NVIDIA GPU clusters and serving live enterprise and government-scale voice traffic. The models are ours, so the whole stack is open to you: you can change the serving engine, the quantization recipe, the kernel, or the model itself, and you will often need to change more than one to hit a target.
The constraints are real and the feedback loop is fast. A latency regression is heard by callers within minutes. A successful optimization shows up directly in infrastructure spend. You will have access to proprietary speech and language data, in-house training and evaluation infrastructure, a close engineering partnership with NVIDIA, and a team that has been building Indian-language AI since 2016.
This role has a clear path to owning the inference platform and its technical direction, or to a deeper specialization in performance and kernel engineering, depending on where your strength lies.
Did you find something suspicious?