Posted on: 23/05/2026
Description :
Role : Agentic AI QE Engineer.
Exp : 4 to 8 years.
Locations : India - Bangalore & Hyderabad.
What You'll Do :
- Design and execute validation strategies for agentic AI systems - from planning and reasoning to action execution and output quality.
- Build and maintain automated eval harnesses to measure accuracy, groundedness, hallucination rates, and safety guardrail adherence across multi-LLM deployments.
- Validate LangChain, RAG, and tool-calling architectures against defined quality benchmarks.
- Develop automated test frameworks in Python/TypeScript using Playwright, API automation, or custom tooling.
- Collaborate with AI engineers to define testable acceptance criteria for agent behaviors.
- Participate in CI/CD integration for automated regression of AI workflow quality.
- Conduct device-level validation (Android, Windows) for on-device AI feature behavior including memory, latency, and battery impact.
Required Skills :
- 4+ Years into Automation Testing.
- 2+ years of credible, hands-on experience in AI, LLM, or agentic AI testing/evaluation.
- Strong understanding of how LLMs work - prompt behavior, context windows, tool use, chain-of-thought.
- Test automation proficiency : Python, TypeScript, or JavaScript; Playwright, Selenium, or Appium.
- Experience building or using eval frameworks (custom harnesses, LLM-as-judge patterns, or similar).
- Excellent written English and ability to work independently without day-to-day guidance.
Nice to Have :
- Experience with LangChain, AutoGen, CrewAI, or similar agent frameworks.
- Device-level testing on Android or Windows hardware.
- Familiarity with Cursor AI or Claude Code for AI-assisted test development.
- Background in MLOps or AI model deployment pipelines.
Did you find something suspicious?