HamburgerMenu
hirist

Lead ML Research Scientist - Speech AI

HR LOBBY
6 - 10 Years
rupee40-60 LPA
Bangalore

Posted on: 05/10/2026

Job Description

Engagement Overview :

We are seeking an experienced Lead ML Research Scientist on a consultancy / contract basis to drive the technical vision and architectural foundation of our next-generation speech AI technology stack.

Our enterprise platform powers voice AI agents for tier-1 institutions in banking, insurance, healthcare, and telecom, handling tens of millions of live production voice minutes. The platform is ISO 27001/27701 and SOC 2 Type II certified, operating across both cloud and on-premise environments.

Existing in-house models include custom STT (tuned for BFSI), language-switch detection, end-of-utterance detection, and real-time noise cancellation running on a high-throughput, low-latency framework. As a Lead Consultant, you will spearhead advanced research and development across Indian languages, code-mixed speech (Hinglish), natural expressive TTS, and full-duplex speech-to-speech foundation models.

Core Responsibilities :

Technical Vision & Strategy :

- Partner with executive leadership to establish multi-quarter technical roadmaps tied directly to production impact - focusing on word error rate (WER), latency, naturalness, compute efficiency, and domain accuracy.

Foundation Model Architecture :

- Design, scale, and pre-train speech foundation models from scratch on large multilingual audio corpora (100k+ hours), guiding architectural choices across self-supervised learning, tokenizers, codecs, and scale.

Research Track Leadership :

- STT : Lead continuous improvements in low-latency streaming ASR, code-switching robustness, contextual biasing (names, domain-specific entities), and domain adaptation for specialized industries.

- TTS : Lead expressive, streaming TTS development for Indian languages - focusing on zero-shot voice cloning, flow-matching, diffusion, neural codec LMs, prosody/emotion control, and mixed-script text parsing.

- Speech-to-Speech (S2S) : Architect full-duplex conversational models utilizing in-house audio codecs, speech-LLM alignment, real-time interruption handling, and turn-taking prediction.

Data Strategy & Infrastructure :

- Formulate data acquisition, synthetic generation, forced alignment, and labeling strategies. Guide multi-node GPU distributed training (100k+ hours, FSDP/DeepSpeed/Megatron) and ensure stable training runs.

Mentorship & Quality Standards :

- Direct and mentor a team of 4 - 8+ speech researchers and engineers; institute rigorous code reviews, experiment tracking, loss-spike debugging, and baseline evaluation standards.

Required Qualifications & Expertise :

- 6 - 10+ years of total experience in Machine Learning, with at least 4+ years dedicated specifically to speech/audio ML (ASR, TTS, neural codecs, or audio-LLMs).

- Education : MS or PhD in CS, Electrical Engineering, or a related field (or equivalent depth demonstrated through shipped production systems or high-impact research).

- Proven Track Record : Proven experience leading foundation model pre-training from scratch.

- Architecture Mastery : Deep hands-on knowledge of Conformer/Zipformer, RNN-T/TDT, CTC, flow-matching/diffusion TTS, neural codecs (EnCodec, DAC, Mimi), and Speech-LLMs.

- Large-Scale Pre-Training : Practical experience with multi-node distributed training infrastructure, scaling laws, compute budgeting, and handling multi-node cluster instabilities.

- Streaming & Telephony Audio : Expertise in building models optimized for streaming inference under low-latency constraints on telephony (8 kHz) or WebRTC audio.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...