Posted on: 31/03/2026
About the Role :
You will design and own scalable speech and audio processing systems including :
- Audio ingestion - preprocessing - tagging - QA - dataset delivery pipelines
- Automated audio quality checks (SNR, clipping, silence detection, noise profiling)
- ASR integration & evaluation (Whisper, Kaldi, SpeechBrain, etc.)
- WER / CER computation and optimization
- Transcription validation workflows
- Audio normalization, resampling, segmentation, chunking
- Structuring enterprise-grade speech datasets
- Speaker verification, dialect tagging, emotion tagging
Candidates must demonstrate hands-on ownership of speech pipeline components.
Required Experience :
- 4 to 7 years of hands-on experience in speech/audio processing
- Strong Python
- Experience with Librosa, torchaudio, PyDub, SpeechBrain, Kaldi or similar
- Experience working with ASR systems in production
- Strong understanding of sampling rates, spectrograms, MFCCs, preprocessing
- Experience computing and improving WER/CER
- Experience building pipelines on AWS / GCP / Azure
Strongly Preferred :
- Strong academic foundation in signal processing, speech technology, or related field preferred
- Experience fine-tuning ASR models
- Speaker diarization
- Multilingual speech systems
- Large-scale dataset QA automation
Not a Fit If :
- Your experience is primarily prompt engineering
- You have only wrapped Whisper APIs without audio-level ownership
- You are a generic full-stack or RAG engineer
Did you find something suspicious?
Posted by
Posted in
Semiconductor/VLSI/EDA
Functional Area
Embedded / Kernel Development
Job Code
1624928