How to test a GenAI assistant before every release
Prompts, models, and retrieval settings change all the time, and most teams judge quality by trying a few questions. A test set written by your experts, run on every change, catches problems before users do.
A GenAI assistant changes more often than most software. Someone rewrites the prompt. The model provider releases a new version. The team adjusts how documents are split for search. Each change can improve some answers and quietly break others.
In many teams, quality is checked by a few people trying a few questions before release. That catches obvious failures. It does not tell you whether the change made last month’s hard cases worse, or whether the assistant now reveals something it should not. Those problems reach users first.
The fix is the one software teams already use for code: a set of tests that runs on every change and holds the release when the results fall short.
Build the test set from real questions
Start with questions people actually ask. Search logs, support tickets, and the questions from a pilot group are all good sources, once personal data has been removed.
For each question, a subject-matter expert writes down what a correct answer must contain and which sources it should cite. That reference answer is the part that cannot be automated, and it is what makes the tests meaningful to the business. Then add edge cases on purpose: questions the assistant should refuse, questions outside its scope, and questions where the honest answer is that it does not know.
Cover more than accuracy
Accuracy is only one thing to test. A useful set also checks grounding, meaning whether the answer comes from the right sources and cites them, and refusal, meaning whether the assistant declines what it should decline. It checks that the assistant stays within what the person asking is allowed to see, and that it ignores instructions hidden in documents or questions that try to change its behavior. It checks for harmful or inappropriate content. And it checks that answers arrive within the time and cost limits you have set.
Some of these failures are tolerable in small numbers and some are not. A slightly weaker answer to an obscure question is a minor problem. A single answer that leaks restricted information is a serious one. Mark the cases where any failure is unacceptable as critical, so they are treated differently when results are scored.
Plant canaries to catch leaks
Some failures are hard to spot in a long answer. A simple technique helps: place unique, made-up strings in restricted documents and in the system prompt. No legitimate answer should ever contain one, so if an answer does, you have found a leak, and a plain text check will catch it every time.
Agree the release gates with the business owner
Before running anything, agree with the owner of the assistant what a passing release looks like. That usually means a minimum pass rate overall and for each category, no failures at all on critical cases such as permissions and injection, limits on average cost and response time, and a cap on how far results may drop compared with the last release.
Write these gates down. They are the assistant’s quality standard, and they give whoever oversees your AI systems a clear basis for approving it.
Run the tests on every change
Connect the tests to your build pipeline so they run whenever a prompt, a model, or a retrieval setting changes. The pipeline scores the change, publishes a report, and blocks the change if a gate fails.
The most useful part of the report is the comparison with the previous release: the list of cases that passed last time and fail now. That list is what the team needs to decide whether a change is worth making.
Add a test for every bug
When a bad answer shows up in production, add it to the test set before you fix it. The fix can then be proved, and the problem cannot come back unnoticed. Refresh the set with new real questions each quarter, so it keeps pace with how people actually use the assistant.
What you get
With tests in place, the team can change prompts and models with confidence, because it knows what each change does. The business owner gets a clear quality standard and a report for every release. And when someone asks whether the assistant is safe to use, there is evidence to show them.
Our GenAI Evaluation Tests provide the test set format, scoring, red-team cases, and release gates for GitHub Actions and Azure DevOps. If you are running or planning an assistant, tell us about it and we will tell you what the first step would be.