Posted on: 05/10/2026
Engagement Overview :
We are seeking an experienced Lead ML Research Scientist on a consultancy / contract basis to drive the technical vision and architectural foundation of our next-generation speech AI technology stack.
Our enterprise platform powers voice AI agents for tier-1 institutions in banking, insurance, healthcare, and telecom, handling tens of millions of live production voice minutes. The platform is ISO 27001/27701 and SOC 2 Type II certified, operating across both cloud and on-premise environments.
Existing in-house models include custom STT (tuned for BFSI), language-switch detection, end-of-utterance detection, and real-time noise cancellation running on a high-throughput, low-latency framework. As a Lead Consultant, you will spearhead advanced research and development across Indian languages, code-mixed speech (Hinglish), natural expressive TTS, and full-duplex speech-to-speech foundation models.
Core Responsibilities :
Technical Vision & Strategy :
- Partner with executive leadership to establish multi-quarter technical roadmaps tied directly to production impact - focusing on word error rate (WER), latency, naturalness, compute efficiency, and domain accuracy.
Foundation Model Architecture :
- Design, scale, and pre-train speech foundation models from scratch on large multilingual audio corpora (100k+ hours), guiding architectural choices across self-supervised learning, tokenizers, codecs, and scale.
Research Track Leadership :
- STT : Lead continuous improvements in low-latency streaming ASR, code-switching robustness, contextual biasing (names, domain-specific entities), and domain adaptation for specialized industries.
- TTS : Lead expressive, streaming TTS development for Indian languages - focusing on zero-shot voice cloning, flow-matching, diffusion, neural codec LMs, prosody/emotion control, and mixed-script text parsing.
- Speech-to-Speech (S2S) : Architect full-duplex conversational models utilizing in-house audio codecs, speech-LLM alignment, real-time interruption handling, and turn-taking prediction.
Data Strategy & Infrastructure :
- Formulate data acquisition, synthetic generation, forced alignment, and labeling strategies. Guide multi-node GPU distributed training (100k+ hours, FSDP/DeepSpeed/Megatron) and ensure stable training runs.
Mentorship & Quality Standards :
- Direct and mentor a team of 4 - 8+ speech researchers and engineers; institute rigorous code reviews, experiment tracking, loss-spike debugging, and baseline evaluation standards.
Required Qualifications & Expertise :
- 6 - 10+ years of total experience in Machine Learning, with at least 4+ years dedicated specifically to speech/audio ML (ASR, TTS, neural codecs, or audio-LLMs).
- Education : MS or PhD in CS, Electrical Engineering, or a related field (or equivalent depth demonstrated through shipped production systems or high-impact research).
- Proven Track Record : Proven experience leading foundation model pre-training from scratch.
- Architecture Mastery : Deep hands-on knowledge of Conformer/Zipformer, RNN-T/TDT, CTC, flow-matching/diffusion TTS, neural codecs (EnCodec, DAC, Mimi), and Speech-LLMs.
- Large-Scale Pre-Training : Practical experience with multi-node distributed training infrastructure, scaling laws, compute budgeting, and handling multi-node cluster instabilities.
- Streaming & Telephony Audio : Expertise in building models optimized for streaming inference under low-latency constraints on telephony (8 kHz) or WebRTC audio.
Did you find something suspicious?