HamburgerMenu
hirist

Job Description

About the job :

Required Skills (AI/ML-1 Level) Technical Skills :

- Strong experience in Generative AI, LLMs, and Agentic AI systems.

- Hands-on expertise with AI evaluation frameworks (RAGAS, DeepEval, TruLens, LangSmith, Promptfoo, etc.).

- Proficiency in Python and AI/ML development libraries.

- Knowledge of Prompt Engineering, prompt testing, and optimization.

- Ability to define and track evaluation metrics such as accuracy, relevance, groundedness, hallucination rate, latency, and user satisfaction.

- Experience in creating automated evaluation pipelines and benchmarking frameworks.

- Strong understanding of AI safety, guardrails, bias testing, and responsible AI practices.

- Familiarity with REST APIs, JSON, vector databases, and knowledge retrieval systems.

- Experience in A/B testing, human-in-the-loop evaluation, and red teaming.

- Strong experience in Manual Testing of AI/GenAI applications, including functional, exploratory, UAT, regression, and end-to-end testing.

- Expertise in validating Agent Reasoning, Tool Calling, Workflow Execution, and Response Quality.

- Hands-on experience in Automation Testing using Python frameworks.

Key Responsibilities :

- Design, execute, and automate evaluation strategies for Agentic AI applications.

- Develop evaluation datasets, test cases, and benchmark suites.

- Measure and improve agent performance, reasoning quality, tool usage, and workflow effectiveness.

- Analyze model outputs and identify hallucinations, biases, safety risks, and failure patterns.

- Collaborate with AI Engineers, Product Teams, and Domain Experts to improve agent quality and reliability.

- Generate evaluation reports, dashboards, and actionable recommendations.

info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...