Posted on: 21/07/2026
Key Responsibilities:
- Design and build scalable data pipelines for collecting, cleaning, labeling, and processing conversational voice data.
- Fine-tune and train speech and language models for improved accuracy, latency, and production performance.
- Develop evaluation datasets, benchmarks, and automated testing pipelines for speech quality and model performance.
- Build human-in-the-loop evaluation workflows to continuously improve model quality.
- Collaborate with engineering teams to deploy models into high-volume, latency-sensitive production environments.
- Monitor production model performance, investigate regressions, and continuously optimize model behavior.
- Improve speech recognition quality across multiple Indian languages, accents, and telephony conditions.
- Optimize inference performance through quantization, distillation, and latency optimization techniques.
Required Skills:
- 3+ years of hands-on Machine Learning experience with production model development.
- Strong programming skills in Python and PyTorch.
- Experience with distributed training and modern fine-tuning techniques such as LoRA, QLoRA, DPO, and RLHF.
- Strong understanding of ML data pipelines including data collection, preprocessing, labeling, deduplication, augmentation, and quality validation.
- Experience designing robust model evaluation frameworks and benchmarking systems.
- Strong understanding of production ML systems and model deployment.
- Excellent analytical and problem-solving skills.
Preferred Skills:
- Experience with Speech AI, ASR, TTS, or Voice AI systems.
- Knowledge of real-time or streaming inference architectures.
- Experience with model optimization techniques including quantization, pruning, and distillation.
- Familiarity with GPU training, distributed computing, and cloud ML infrastructure.
- Experience deploying ML models at scale.
Did you find something suspicious?