Posted on: 19/08/2026
Role Overview:
As an AI Quality & Evaluation Engineer, you will be at the forefront of ensuring the reliability, safety, and performance of our generative AI and machine learning deployments. You will work closely with data scientists, product managers, and client stakeholders to design rigorous evaluation frameworks that bridge the gap between model output and real-world utility. Your work directly influences the quality of AI-driven decisions, ensuring that our solutions remain robust, unbiased, and aligned with client-specific business objectives.
Key Responsibilities:
- Design and implement automated evaluation pipelines to benchmark model performance against defined accuracy, latency, and safety metrics.
- Develop comprehensive test datasets and golden sets to identify edge cases and potential failure modes in large language models and predictive systems.
- Collaborate with cross-functional teams to translate complex business requirements into measurable quality assurance criteria for AI models.
- Conduct root-cause analysis on model hallucinations or performance regressions to provide actionable insights for model fine-tuning and optimization.
- Establish human-in-the-loop feedback mechanisms to refine model outputs and ensure alignment with industry-specific compliance and quality standards.
Required Skillset:
- Demonstrated proficiency in Python and familiarity with machine learning frameworks such as PyTorch or TensorFlow for building evaluation scripts.
- Strong understanding of LLM evaluation methodologies, including RAG evaluation, prompt engineering, and metrics like BLEU, ROUGE, or semantic similarity.
- Ability to communicate complex technical findings to non-technical stakeholders, ensuring transparency in model performance and limitations.
- Analytical mindset with the ability to design structured experiments and interpret statistical data to drive product improvements.
- A degree in Computer Science, Data Science, or a related quantitative field, reflecting a strong foundation in algorithmic thinking.
- Adaptability to work in a hybrid environment, collaborating effectively with distributed teams to meet project milestones in a fast-paced consulting setting.
Did you find something suspicious?