Posted on: 09/10/2026
LLM/AI Model Engineer - Training, Evaluation & Benchmarking
Location : New Delhi, India (On-site)
Employment Type : Full-time
Experience : 3 - 5 years
Industry : Enterprise AI / Generative AI / Machine Learning / NLP
About QubeLabs :
QubeLabs is building the next generation of Enterprise AI Systems that transform workforce operations and intelligence for the financial services industry.
Our platform combines Conversational AI, Agentic AI, workflow automation, proprietary language models and enterprise intelligence. Our flagship product, QubeLabs Workmate, serves banks, NBFCs, MFIs, wealth management firms, insurance companies and fintechs across India and Europe.
At the core of our AI platform is the Vectro Series, our proprietary family of language models designed for enterprise and financial-services use cases.
We are looking for an LLM/AI Model Engineer to build the training pipelines, datasets and evaluation infrastructure required to continuously improve the Vectro Series.
Role Overview :
You will work closely with the Head of AI, Research Partners and engineering teams to implement and operationalize LLM training approaches and build a reliable model improvement lifecycle: Data Preparation - Training - Evaluation - Benchmarking - Error Analysis - Model Improvement.
The role combines hands-on LLM fine-tuning, dataset engineering, tokenization, evaluation framework development and model quality management.
Key Responsibilities :
1. LLM Training & Fine-tuning :
- Build and maintain pipelines for training and fine-tuning open-source foundation models for the Vectro Series.
- Implement supervised fine-tuning (SFT), parameter-efficient fine-tuning (PEFT), LoRA/QLoRA and domain adaptation techniques; support continued pre-training where required.
- Manage training configurations, checkpoints, model versions and experiment tracking.
- Optimize training workflows for computational efficiency, reproducibility and reliability.
- Translate model development approaches defined with the Head of AI and Research Partners into practical, scalable training pipelines.
2. Training Data & Tokenization :
- Build data pipelines to collect, clean, filter, normalize, deduplicate and validate training data.
- Develop instruction-tuning, SFT and domain-specific datasets for financial-services and enterprise use cases.
- Implement tokenization workflows, tokenizer configuration, vocabulary management, sequence packing, truncation and padding.
- Address data representation challenges involving financial terminology, numerical values and multilingual content.
- Maintain dataset versioning, lineage and quality controls, incorporating identified data gaps and relevant model failures into future training data.
3. Golden Dataset & Model Evaluation :
- Build and maintain QubeLabs' Golden Dataset to evaluate key model capabilities and target use cases.
- Develop automated evaluation pipelines and reusable evaluation harnesses, combining automated metrics with human assessment where appropriate.
- Evaluate accuracy, reasoning, factuality, instruction following, safety, multilingual performance and financial-domain capabilities.
- Establish consistent evaluation methodology and regression tests to measure changes across model versions.
4. Benchmarking & Error Analysis :
- Evaluate Vectro against relevant public benchmarks and proprietary QubeLabs BFSI/enterprise benchmarks.
- Compare results with relevant open-source and commercial models using consistent evaluation conditions.
- Analyze failures, identify root causes across data, training and model behavior, and translate findings into actionable improvements.
- Maintain reproducible benchmark results and reports to track model strengths, limitations and progress.
5. Model Quality & Release :
- Define model release quality gates and ensure each major release meets agreed evaluation and benchmark criteria.
- Support red-teaming, robustness and safety testing.
- Coordinate with the Head of AI, Research Partners and engineering teams to ensure validated model improvements are suitable for integration into production systems.
Experience :
- 3 - 5 years in Machine Learning, Deep Learning, NLP, Generative AI or related fields.
- Hands-on experience training or fine-tuning LLMs using open-source foundation models.
- Practical experience building training pipelines, preparing datasets and evaluating model performance.
- Experience with GPU-based training; distributed training is an advantage.
- Financial-services, multilingual AI or enterprise AI experience is preferred.
Required Skills :
- LLM & Machine Learning : Large Language Models, Transformer architectures, Generative AI, NLP, Deep Learning, SFT, PEFT, LoRA/QLoRA, fine-tuning and continued pre-training.
- Frameworks & Infrastructure : Python, PyTorch, Hugging Face Transformers, Hugging Face Tokenizers, GPU computing, experiment tracking and ML pipelines.
- Data & Tokenization : Tokenization, tokenizer configuration, text preprocessing, sequence packing, dataset construction, data quality, versioning and lineage.
- Evaluation & Benchmarking : Golden datasets, LLM evaluation, evaluation harnesses, automated and human evaluation, LLM-as-a-Judge, public and domain-specific benchmarks, regression testing and error analysis.
The role requires the ability to build both the model training pipeline and the measurement system that determines whether the model has actually improved.
What Success Looks Like :
- Reliable and reproducible training and fine-tuning pipelines for the Vectro Series.
- High-quality training datasets, tokenization workflows and a comprehensive Golden Dataset.
- Automated evaluation and benchmarking infrastructure with measurable model quality gates.
- Clear, reproducible evidence of model performance against public, proprietary BFSI and relevant competing-model benchmarks.
- A continuous improvement loop that converts evaluation findings into measurable gains across successive Vectro releases.
Key Performance Indicators :
- Training pipeline reliability and experiment throughput.
- Training data quality and coverage.
- Golden Dataset and evaluation coverage.
- Benchmark reproducibility and model performance improvement.
- Accuracy of error diagnosis and effectiveness of regression detection.
- Time required to complete training, evaluation and benchmarking cycles.
- Compliance with model release quality criteria.
Why Join QubeLabs?
- Build the proprietary Vectro Series of enterprise and financial-services language models.
- Develop model training, evaluation and benchmarking infrastructure from the ground up.
- Work closely with the Head of AI and Research Partners.
- Solve real-world AI challenges across financial services.
- Develop proprietary datasets, benchmarks and model improvement systems.
- Contribute directly to production AI systems serving customers across India and Europe.
Preferred Candidate Profile :
- B.Tech, M.Tech, MS or PhD in Computer Science, AI, ML, Mathematics or a related field.
- 3 - 5 years of relevant experience with strong practical skills in Python, PyTorch, LLM fine-tuning and training data pipelines.
- Working knowledge of tokenization, evaluation datasets, benchmarking and regression testing.
- Experience with open-source models such as Llama, Qwen, Mistral or similar is preferred.
- Ability to independently build reliable model training and evaluation systems from scratch.
QubeLabs is an equal opportunity employer and values diversity and inclusion. We encourage applications from qualified candidates regardless of background or identity.
Did you find something suspicious?