Posted on: 27/08/2026
Role Overview: Build AI Systems We Can Trust
We are seeking a Senior AI Evaluation & Reliability Engineer passionate about solving one of the most critical challenges in modern software: How do we guarantee an AI system is accurate, improving, and delivering tangible business value?
In this role, you will treat model performance and agentic reliability as core software engineering disciplines. Moving far beyond manual prompt inspection, you will design production-grade evaluation systems for LLMs, RAG applications, and multi-agent workflows. Beyond the code, you will serve as a trusted technical advisor to our North American enterprise clients, a mentor to our engineering teams, and a key driver connecting AI performance directly to ROI.
What You Will Own :
1. Architect Production-Grade Evaluation Pipelines :
- Build and maintain end-to-end programmatic evaluation frameworks to continuously validate production LLM outputs, RAG architectures, and multi-agent systems.
- Integrate automated evaluation suites into CI/CD workflows (e.g., GitHub Actions, GitLab CI) to catch regressions before deployment.
- Establish statistically rigorous metric suites measuring Faithfulness, Context Precision, Answer Relevance, Semantic Drift, and Task Completion Rates.
2. Engineer "LLM-as-a-Judge" Systems & Manage Economics :
- Design and optimize scalable LLM-as-a-Judge frameworks, building calibration loops against human ground-truth datasets to mitigate judge-specific biases.
- Manage Evaluation ROI: Running comprehensive evaluations burns tokens. You will design tiered evaluation routing using fast heuristic graders for simple checks and powerful LLMs for complex judging, balancing quality, latency, and inference spend.
3. Benchmark RAG & Synthetic Data :
- Systematically benchmark RAG systems across embedding models, vector indexing, chunking strategies, and re-ranking algorithms.
- Generate programmatic synthetic datasets and adversarial scenarios to stress-test agent reasoning, tool-calling precision, and failure recovery.
4. Drive Enterprise Consulting & Presales :
- Act as the principal technical authority for North American clients, leading architectural discussions and translating complex probabilistic metrics into actionable business strategies.
- Build executive-friendly AI quality and ROI scorecards to help clients establish Service Level Objectives (SLOs) and risk-acceptance thresholds.
- Actively partner with Sales and Product teams in discovery calls, shaping technical proposals and demonstrating our AI capabilities to win strategic engagements.
5. Coach Engineers & Raise the Bar :
- Mentor internal teams, moving engineers from basic prompt experimentation to disciplined, programmatic AI engineering.
- Build reusable frameworks, playbooks, and internal accelerators to scale our AI evaluation capabilities organization-wide.
What We Are Looking For :
AI Evaluation Mastery :
- Proven track record of building and deploying programmatic evaluation systems in production.
- Deep, hands-on experience with LLM-as-a-Judge calibration, RAG evaluation, and mitigating evaluation bias.
- Mastery of modern evaluation tooling: DeepEval, Ragas, TruLens, LangSmith, DSPy, Promptfoo, or equivalent.
Software Engineering & Backend Foundations :
- Programming Depth: Advanced proficiency in Python and/or TypeScript.
- Architecture: Strong experience with asynchronous API design, CI/CD, and vector databases (e.g., Pinecone, Qdrant, Milvus, Chroma).
- Cost Optimization: Proven ability to manage strict cost-control strategies for model inference and evaluation loops.
Consulting & Leadership :
- Exceptional verbal and written communication skills, with a history of interfacing directly with enterprise stakeholders (CTO/COO level).
- A consultative problem-solving mindset: the ability to clearly articulate the trade-offs between model latency, inference cost, and output accuracy.
- A strong desire to mentor, lead, and stay at the bleeding edge of emerging AI technologies.
Why This Role Matters :
The next generation of industry-leading AI companies won't just be the ones with access to the best models. They will be the ones that can confidently measure AI, trust AI, and translate that reliability into actual business value. If you want to build that future with us, we'd love to hear from you.
Did you find something suspicious?