Posted on: 15/09/2026
Role Overview :
We are looking for a Lead Data Scientist to own and drive AI/ML initiatives across NLP, Traditional ML, and Generative AI. The role will focus heavily on data problem formulation, training data strategy, model training, and building production-grade AI systems that continuously improve from real-world outcomes.
You will work closely with Data Engineering, GenAI/NLP, and Product teams to build scalable AI solutions and take ownership of technical architecture, model development, and AI/ML best practices.
What You'll Do :
- Own data problem formulation for AI systems, including structuring, labeling, and mining historical decision data and domain signals.
- Lead training data strategies by building high-quality labeled datasets from real-world data and prioritizing annotation efforts.
- Drive large-scale data mining across high-cardinality datasets to identify rare but valuable patterns.
- Own the architecture and evolution of AI platforms covering real-time inference, batch processing, and model training pipelines.
- Build multi-modal AI systems combining unstructured text, structured data, and domain knowledge.
- Lead technical evaluation of AI tools, platforms, and vendors, including build-vs-buy decisions.
- Mentor AI engineers across NLP, Traditional ML, and GenAI and lead architecture reviews.
- Collaborate with Data Engineering, GenAI/NLP, and Product teams to align data and model training strategies with product requirements.
What We're Looking For :
- 6+ years of experience building and deploying production AI/ML systems, with strong emphasis on solving data problems.
- Deep experience working with complex, domain-specific data at scale.
- Strong expertise across multiple AI domains, including NLP, Traditional ML, and GenAI.
- Expert-level proficiency in PyTorch or TensorFlow, with hands-on experience in model training at scale.
- Proven experience where data quality, labeling strategy, and training data curation were key challenges.
Did you find something suspicious?