Posted on: 25/09/2026
Role : Senior/Staff AI Evaluation Engineer
About the Role :
We are looking for a Senior/Staff AI Evaluation Engineer to lead the benchmarking, validation, reliability, and safety evaluation of next-generation Agentic AI platforms. This role sits at the intersection of Quality Engineering, Software Engineering, and AI/ML.
What You'll Do :
- Design and implement automated evaluation frameworks for LLM, RAG, and agent-based applications.
- Build and maintain golden datasets, benchmark datasets, test datasets, and regression suites for AI evaluation.
- Develop measurable evaluation criteria for accuracy, relevance, consistency, safety, reliability, and agent behavior.
- Perform adversarial testing to identify hallucinations, prompt vulnerabilities, unsafe behavior, and edge cases.
- Evaluate LangGraph-based and multi-agent systems at the node, state-transition, routing, and workflow levels.
- Use LangSmith for tracing, debugging, experiments, datasets, and evaluation workflows.
- Work extensively with LangChain and LangGraph, including sub-graphs and conditional routing.
- Build Python-based automation for functional, regression, integration, and end-to-end AI testing.
- Evaluate prompts, embeddings, vector search, RAG pipelines, tool calling, and agentic workflows.
- Validate AI systems against safety, security, privacy, and Responsible AI expectations.
- Integrate evaluation and regression testing into GitHub-based CI/CD workflows.
- Work with AWS and Amazon Bedrock to validate AI-powered application architectures.
- Perform API and backend validation using REST APIs, JSON, SQL, and modern application architectures.
- Collaborate with Engineering, Product, Data, and Quality teams to investigate failures and drive improvements.
- Take ownership of ambiguous AI quality problems and convert them into repeatable, measurable evaluation approaches.
What We're Looking For :
- 6+ years of experience in software engineering, quality engineering, test automation, AI engineering, ML engineering, or a related technical discipline.
- Demonstrable hands-on experience testing or evaluating LLM, NLP, ML, RAG, or agentic AI applications.
- Strong Python development and automation experience.
- Deep hands-on knowledge of LangChain, LangGraph, and LangSmith.
- Experience evaluating LangGraph node execution, state transitions, conditional routing, sub-graphs, or multi-agent workflows.
- Strong understanding of LLMs, prompt engineering, embeddings, vector databases/search, RAG, and AI agents.
- Experience building golden datasets, benchmark datasets, or structured AI test datasets.
- Experience with adversarial testing and AI safety evaluation.
- Practical understanding of security, privacy, and Responsible AI evaluation.
- Strong GitHub experience covering repositories, branching, pull requests, code reviews, and CI/CD.
- Experience with Claude Code or comparable AI-assisted development tools.
- Experience with AWS and Amazon Bedrock.
- Strong REST API, JSON, and SQL knowledge.
- Strong experience in automated, regression, integration, and end-to-end testing.
- Excellent analytical, troubleshooting, communication, and problem-solving abilities.
The job is for:
Did you find something suspicious?