Posted on: 03/06/2026
Role : Senior AI Testing Engineer (Generative AI)
Location : Bangalore
Experience : 58 Years
Role Overview :
We are looking for a Senior AI Testing Engineer to own quality across our Generative AI products and platform.
This role is fundamentally about engineering quality into AI systems not running test scripts. The ideal candidate will design evaluation frameworks, build automated testing pipelines, and define what good looks like for LLM outputs, RAG systems, AI agents, and voice AI applications.
You will work closely with AI engineers and product teams to ensure AI systems are reliable, safe, and continuously improving over time. If you understand how LLMs fail, know how to identify hallucinations, and want to build quality infrastructure for production AI at scale, this role is for you.
Key Responsibilities :
Evaluation Strategy & Frameworks :
Design and own comprehensive testing strategies for Generative AI products, including :
- LLM applications
- RAG pipelines
- AI agents
- Voice AI systems
- Workflow automation
Define evaluation methodologies covering :
- Functional testing
- Response quality assessment
- Hallucination detection
- Safety and guardrail testing
- Prompt injection testing
- Bias and toxicity testing
- Retrieval quality evaluation
- Latency benchmarking
- Agent workflow validation
Build reusable AI testing frameworks and automation pipelines for continuous evaluation.
Create datasets, benchmark suites, and golden test sets for GenAI evaluation.
Automated Evaluation :
- Develop automated evaluation pipelines using LLM-as-a-Judge and hybrid evaluation methods.
- Implement CI/CD-integrated AI evaluation pipelines.
- Drive observability and monitoring strategies for production AI systems.
Quality Standards & Collaboration :
- Define measurable quality KPIs for AI systems.
- Establish testing standards, best practices, and governance processes for GenAI applications.
- Collaborate closely with AI engineers, product, and platform teams to embed quality throughout the development lifecycle.
Required Skills & Experience :
Testing & Engineering Experience :
- 5 to 8 years of experience in Software Testing, QA Engineering, SDET, or Test Automation.
- 2 to 3 years of hands-on experience testing or evaluating production-grade Generative AI or LLM-based systems.
- Strong test automation skills in Python.
- Experience designing scalable automated testing frameworks.
- Familiarity with API testing, integration testing, and performance testing.
Generative AI Knowledge :
- Solid understanding of LLM systems and common failure modes.
- Experience with RAG architectures, prompt engineering, AI agents, embedding models, and vector databases.
- Understanding of LLM evaluation methodologies and AI system failure patterns.
GenAI Testing Frameworks :
Hands-on experience with one or more GenAI evaluation frameworks, such as :
- DeepEval
- Ragas
- LangSmith
- Promptfoo
- TruLens
- OpenAI Evals
- LangChain Evaluation Tools
Quality Engineering :
Expertise in :
- Test strategy & planning
- Test automation architecture
- Defect lifecycle management
- Quality metrics
- Ability to define and track measurable quality KPIs for AI systems.
Preferred Qualifications :
- Experience with Cloud Platforms (AWS, Azure, or GCP).
- Familiarity with MLOps / LLMOps workflows.
- Experience with CI/CD pipelines and DevOps practices.
- Exposure to monitoring and observability tooling for AI systems.
- Understanding of security and compliance for GenAI products.
- Experience with Conversational AI or Voice AI systems.
Did you find something suspicious?
Posted by
Abhinayani N
HR Officer at National Entrepreneurship Network
Last Active: NA as recruiter has posted this job through third party tool.
Posted in
Quality Assurance
Functional Area
QA & Testing
Job Code
1641455