HamburgerMenu
hirist

Senior/Staff AI Evaluation Engineer - Python

Haparz
6 - 9 Years
Multiple Locations

Posted on: 25/09/2026

Job Description

Role : Senior/Staff AI Evaluation Engineer

About the Role :

We are looking for a Senior/Staff AI Evaluation Engineer to lead the benchmarking, validation, reliability, and safety evaluation of next-generation Agentic AI platforms. This role sits at the intersection of Quality Engineering, Software Engineering, and AI/ML.

What You'll Do :

- Design and implement automated evaluation frameworks for LLM, RAG, and agent-based applications.

- Build and maintain golden datasets, benchmark datasets, test datasets, and regression suites for AI evaluation.

- Develop measurable evaluation criteria for accuracy, relevance, consistency, safety, reliability, and agent behavior.

- Perform adversarial testing to identify hallucinations, prompt vulnerabilities, unsafe behavior, and edge cases.

- Evaluate LangGraph-based and multi-agent systems at the node, state-transition, routing, and workflow levels.

- Use LangSmith for tracing, debugging, experiments, datasets, and evaluation workflows.

- Work extensively with LangChain and LangGraph, including sub-graphs and conditional routing.

- Build Python-based automation for functional, regression, integration, and end-to-end AI testing.

- Evaluate prompts, embeddings, vector search, RAG pipelines, tool calling, and agentic workflows.

- Validate AI systems against safety, security, privacy, and Responsible AI expectations.

- Integrate evaluation and regression testing into GitHub-based CI/CD workflows.

- Work with AWS and Amazon Bedrock to validate AI-powered application architectures.

- Perform API and backend validation using REST APIs, JSON, SQL, and modern application architectures.

- Collaborate with Engineering, Product, Data, and Quality teams to investigate failures and drive improvements.

- Take ownership of ambiguous AI quality problems and convert them into repeatable, measurable evaluation approaches.

What We're Looking For :

- 6+ years of experience in software engineering, quality engineering, test automation, AI engineering, ML engineering, or a related technical discipline.

- Demonstrable hands-on experience testing or evaluating LLM, NLP, ML, RAG, or agentic AI applications.

- Strong Python development and automation experience.

- Deep hands-on knowledge of LangChain, LangGraph, and LangSmith.

- Experience evaluating LangGraph node execution, state transitions, conditional routing, sub-graphs, or multi-agent workflows.

- Strong understanding of LLMs, prompt engineering, embeddings, vector databases/search, RAG, and AI agents.

- Experience building golden datasets, benchmark datasets, or structured AI test datasets.

- Experience with adversarial testing and AI safety evaluation.

- Practical understanding of security, privacy, and Responsible AI evaluation.

- Strong GitHub experience covering repositories, branching, pull requests, code reviews, and CI/CD.

- Experience with Claude Code or comparable AI-assisted development tools.

- Experience with AWS and Amazon Bedrock.

- Strong REST API, JSON, and SQL knowledge.

- Strong experience in automated, regression, integration, and end-to-end testing.

- Excellent analytical, troubleshooting, communication, and problem-solving abilities.

The job is for:

May work from home
info-icon

Did you find something suspicious?

Similar jobs that you might be interested in

Loading chat...