Posted on: 09/09/2026
About the job :
Required Skills (AI/ML-1 Level) Technical Skills :
- Strong experience in Generative AI, LLMs, and Agentic AI systems.
- Hands-on expertise with AI evaluation frameworks (RAGAS, DeepEval, TruLens, LangSmith, Promptfoo, etc.).
- Proficiency in Python and AI/ML development libraries.
- Knowledge of Prompt Engineering, prompt testing, and optimization.
- Ability to define and track evaluation metrics such as accuracy, relevance, groundedness, hallucination rate, latency, and user satisfaction.
- Experience in creating automated evaluation pipelines and benchmarking frameworks.
- Strong understanding of AI safety, guardrails, bias testing, and responsible AI practices.
- Familiarity with REST APIs, JSON, vector databases, and knowledge retrieval systems.
- Experience in A/B testing, human-in-the-loop evaluation, and red teaming.
- Strong experience in Manual Testing of AI/GenAI applications, including functional, exploratory, UAT, regression, and end-to-end testing.
- Expertise in validating Agent Reasoning, Tool Calling, Workflow Execution, and Response Quality.
- Hands-on experience in Automation Testing using Python frameworks.
Key Responsibilities :
- Design, execute, and automate evaluation strategies for Agentic AI applications.
- Develop evaluation datasets, test cases, and benchmark suites.
- Measure and improve agent performance, reasoning quality, tool usage, and workflow effectiveness.
- Analyze model outputs and identify hallucinations, biases, safety risks, and failure patterns.
- Collaborate with AI Engineers, Product Teams, and Domain Experts to improve agent quality and reliability.
- Generate evaluation reports, dashboards, and actionable recommendations.
Did you find something suspicious?
Posted by
Sumandeep Tuteja
Last Active: NA as recruiter has posted this job through third party tool.
Posted in
Quality Assurance
Functional Area
QA & Testing
Job Code
1669809