Posted on: 01/10/2026
About the Role:
You will own a research direction (STT, TTS or speech-to-speech) end to end. That means designing new architectures, pre-training foundation models from scratch on large multilingual audio, and getting them into production, while mentoring engineers on the team.
What you will do:
- Lead one research track and set its roadmap with the CTO and product leads, tied to customer metrics (accuracy, naturalness, latency, cost per minute).
- Design novel architectures and pre-train speech foundation models from scratch on large multilingual corpora (100k+ hours), including self-supervised pre-training, tokenizer and codec design, and scaling experiments. Fine-tuning and continued training are tools, not the whole job.
- STT: streaming ASR with low time-to-first-token, code-switch robustness, contextual biasing for names and entities, domain adaptation for BFSI and healthcare.
- TTS: expressive, low-latency streaming TTS for Indian languages; voice cloning, prosody and emotion control, pronunciation of numbers, dates and mixed-script text.
- Speech-to-speech: Full-duplex conversational models built in-house: our own audio tokenizers and codecs, speech-LLM pre-training and alignment, interruption and turn-taking, end-of-utterance prediction.
- Build the data strategy: sourcing, licensing, synthetic data generation, labelling quality and consent-compliant use of production audio.
- Define evaluation standards and benchmarks that predict real call outcomes, and run honest comparisons against external vendors.
- Make serving trade-offs with the platform team: distillation, quantisation, GPU concurrency per model, on-prem footprints.
- Mentor 1 - 3 engineers; review experiment designs and code.
- Publish or open-source selectively where it strengthens our position.
What you bring:
- 3 - 6 years in ML, with at least 2 - 3 years focused on speech (ASR, TTS, speaker/voice or audio-language models).
- MS or PhD in CS, EE or related field, or equivalent depth shown through shipped work.
- Track record of pre-training at least one speech or audio model from scratch (not only fine-tuning) that reached production or strong published results.
- Deep knowledge of modern architectures: Conformer/Zipformer, RNN-T/TDT, encoder-decoder ASR, flow-matching and diffusion TTS, neural codecs (EnCodec, DAC, Mimi), speech-LLMs.
- Large-scale pre-training experience: distributed training (FSDP/DeepSpeed/Megatron) on multi-node GPU clusters, 100k+ audio hours, scaling laws and compute budgeting, data loading at scale, debugging loss spikes and unstable runs.
- Strong experimental design: ablations, significance, avoiding test-set leakage.
- Clear written and spoken communication with engineers, product and customers.
Nice to have:
- Publications at Interspeech, ICASSP, ACL/EMNLP, NeurIPS or similar.
- Experience with Indic languages, low-resource languages or code-switching.
- Built streaming or real-time speech systems with strict latency SLAs.
- Experience with RL or preference tuning for speech (e.g. naturalness rewards) or LLM post-training.
- Contributions to open-source speech toolkits.
What success looks like in 12 months :
- Your track has shipped a model pre-trained from scratch by Blue Machines that beats the vendor it replaces on our benchmark and reduces cost or latency.
- A clear, reproducible training and evaluation stack for your area.
- Engineers you mentor are running experiments independently.
Did you find something suspicious?