Posted on: 19/08/2026
Job Description :
Experience :
7-10 years in ML/speech engineering, 2+ years architecting production voice pipelines.
Key Responsibilities :
- Own end-to-end architecture : ASR ? NLU/LLM ? dialogue manager ? TTS ? telephony
- Make build-vs-buy calls on ASR/TTS vendors (Deepgram, Azure Speech, ElevenLabs, or open-source Whisper/Coqui)
- Own latency budget across the pipeline (target sub-1.5s round-trip for live calls)
- Mentor Conversational AI and ML/Voice engineers; own technical hiring bar
Must-Have Technical Qualifications :
- Production experience with at least one ASR engine (Whisper, Deepgram, Google STT, Azure Speech) at scale
- Hands-on experience with LLM-based dialogue systems (function calling, RAG, or fine-tuning)
- Strong grasp of telephony integration : SIP, WebRTC, Twilio/Exotel/Ozonetel or similar
- Experience optimizing for latency and cost simultaneously in a live voice pipeline
- Comfortable working across Hindi/Hinglish and English accent handling for Indian/UAE markets
- Track record of hitting production accuracy benchmarks : 95%+ transcription (WER-based) accuracy and 90%+ intent classification accuracy at scale not just in a demo
- Experience handling dialect variation and code-switching (e.g. Hinglish, Gulf Arabic variants) in production ASR/NLU
- Successfully deployed and managed Voicebots in a production environment.
- Experience handling 5000+ live customer calls per day.
Good-to-Have :
- Prior experience at Amazon Alexa, Microsoft Cortana/Speech, Google Assistant, or a voice-AI startup
- Exposure to Arabic ASR/TTS for UAE market
- 2+ years hands-on with LLMs/Generative AI in a production (not POC) setting.
Did you find something suspicious?