HamburgerMenu
hirist

ML Research Engineer - Speech/Audio

Eliteeye Consulting
3 - 6 Years
Bangalore

Posted on: 01/10/2026

Job Description

About the Role:

You will own a research direction (STT, TTS or speech-to-speech) end to end. That means designing new architectures, pre-training foundation models from scratch on large multilingual audio, and getting them into production, while mentoring engineers on the team.

What you will do:

- Lead one research track and set its roadmap with the CTO and product leads, tied to customer metrics (accuracy, naturalness, latency, cost per minute).

- Design novel architectures and pre-train speech foundation models from scratch on large multilingual corpora (100k+ hours), including self-supervised pre-training, tokenizer and codec design, and scaling experiments. Fine-tuning and continued training are tools, not the whole job.

- STT: streaming ASR with low time-to-first-token, code-switch robustness, contextual biasing for names and entities, domain adaptation for BFSI and healthcare.

- TTS: expressive, low-latency streaming TTS for Indian languages; voice cloning, prosody and emotion control, pronunciation of numbers, dates and mixed-script text.

- Speech-to-speech: Full-duplex conversational models built in-house: our own audio tokenizers and codecs, speech-LLM pre-training and alignment, interruption and turn-taking, end-of-utterance prediction.

- Build the data strategy: sourcing, licensing, synthetic data generation, labelling quality and consent-compliant use of production audio.

- Define evaluation standards and benchmarks that predict real call outcomes, and run honest comparisons against external vendors.

- Make serving trade-offs with the platform team: distillation, quantisation, GPU concurrency per model, on-prem footprints.

- Mentor 1 - 3 engineers; review experiment designs and code.

- Publish or open-source selectively where it strengthens our position.

What you bring:

- 3 - 6 years in ML, with at least 2 - 3 years focused on speech (ASR, TTS, speaker/voice or audio-language models).

- MS or PhD in CS, EE or related field, or equivalent depth shown through shipped work.

- Track record of pre-training at least one speech or audio model from scratch (not only fine-tuning) that reached production or strong published results.

- Deep knowledge of modern architectures: Conformer/Zipformer, RNN-T/TDT, encoder-decoder ASR, flow-matching and diffusion TTS, neural codecs (EnCodec, DAC, Mimi), speech-LLMs.

- Large-scale pre-training experience: distributed training (FSDP/DeepSpeed/Megatron) on multi-node GPU clusters, 100k+ audio hours, scaling laws and compute budgeting, data loading at scale, debugging loss spikes and unstable runs.

- Strong experimental design: ablations, significance, avoiding test-set leakage.

- Clear written and spoken communication with engineers, product and customers.

Nice to have:

- Publications at Interspeech, ICASSP, ACL/EMNLP, NeurIPS or similar.

- Experience with Indic languages, low-resource languages or code-switching.

- Built streaming or real-time speech systems with strict latency SLAs.

- Experience with RL or preference tuning for speech (e.g. naturalness rewards) or LLM post-training.

- Contributions to open-source speech toolkits.

What success looks like in 12 months :

- Your track has shipped a model pre-trained from scratch by Blue Machines that beats the vendor it replaces on our benchmark and reduces cost or latency.

- A clear, reproducible training and evaluation stack for your area.

- Engineers you mentor are running experiments independently.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...