Posted on: 19/08/2026
Role Overview :
As an AI Quality & Evaluation Engineer, you will be at the forefront of ensuring the reliability, safety, and performance of our next-generation generative AI systems. You will bridge the gap between raw model outputs and production-grade intelligence by designing rigorous evaluation frameworks that measure accuracy, bias, and alignment. Working closely with data scientists, product managers, and software engineers, you will translate complex model behaviors into actionable quality metrics. Your work directly influences the end-user experience, ensuring that our AI solutions remain trustworthy, scalable, and highly effective in solving real-world business challenges.
Key Responsibilities :
- Design and implement automated evaluation pipelines to benchmark LLM performance against human-centric quality standards.
- Develop custom test suites and synthetic datasets to stress-test models, ensuring robustness across diverse edge cases and user scenarios.
- Collaborate with cross-functional teams to define success metrics for AI features, driving continuous improvement in model precision and recall.
- Analyze model failure modes and provide data-driven insights to research teams to accelerate the iterative refinement process.
- Build scalable automation frameworks that integrate seamlessly into our CI/CD pipelines, reducing the time-to-market for high-quality AI deployments.
Required Skillset :
- Demonstrated proficiency in Python for building evaluation scripts, data processing, and automation testing frameworks.
- Deep understanding of LLM architectures and the nuances of evaluating generative models, including RAG systems and fine-tuned agents.
- Proven ability to translate ambiguous quality requirements into structured, measurable testing protocols.
- Strong analytical mindset with the capacity to communicate complex technical findings to non-technical stakeholders clearly and concisely.
- Experience working in a fast-paced, collaborative environment, with the flexibility to thrive in our Bangalore-based hybrid office setup.
- A degree in Computer Science, Data Science, or a related quantitative field, supported by 3 - 6 years of hands-on experience in AI/ML testing and quality assurance.
The job is for:
Did you find something suspicious?