Posted on: 16/07/2026
What You'll Do :
- Design and run training and fine-tuning pipelines for large vision-language models on GPU clusters.
- Build multimodal data pipelines ingestion, filtering, deduplication, synthetic generation, and quality assurance.
- Implement and experiment with new architectures and training techniques from research.
- Build evaluation harnesses, benchmarks, and automated regression tracking.
- Optimise models for inference quantisation, batching, and serving infrastructure.
- Build robust pipelines and integrations that put vision model capabilities in the hands of end users.
- Translate real-world problems into well-scoped ML tasks with the right data and evaluation strategy.
- Work directly with clients to understand their use cases document processing, visual search, form extraction and own the solution end to end.
- Build production-grade systems on top of Sarvam Vision and open-source models: multimodal pipelines, retrieval-augmented workflows, and structured output extraction.
- Debug and improve deployed solutions latency, accuracy, edge cases, and integration with client infrastructure.
What We're Looking For :
- Strong Python and PyTorch comfortable reading and modifying model internals.
- Hands-on experience training or fine-tuning large models, including debugging broken runs.
- Experience building data pipelines at scale.
- Solid grounding in transformer architectures and modern training techniques.
- Comfort with ambiguity the roadmap is not fully pre-specified.
- Strong focus on secure coding practices, code quality, and system reliability.
Did you find something suspicious?