Reference test set
Real questions with reference answers written by your subject-matter experts, plus edge cases.
Tests that measure assistant and agent quality before and after launch.
A test framework for assistants, copilots, and agents. We build a set of questions and expected answers from your own content with your subject-matter experts. We add edge cases and adversarial prompts. Every change to a prompt, model, or retrieval setting is scored for accuracy, use of sources, safety, and cost, and releases wait on the results.
Download the PDFTest results for every release of your generative AI systems.
The problem it solves, what is included, how it works, the technical components, and how we adapt it with you.
Your download has started.
We also emailed the link to . It stays valid for 7 days.
Download didn't start? Get the PDF
Want to see how it would fit your data? Talk to us.
Assistants change constantly: new prompts, new models, new retrieval settings. Without a test set, quality is judged by a few people trying a few questions, and regressions reach users first.
GenAI Evaluation Tests score every change against questions your experts wrote, and hold the release when results fall short.
Four parts, each adapted to your data, platforms, and controls. What we adapt for you is yours to keep.
Real questions with reference answers written by your subject-matter experts, plus edge cases.
Accuracy, use of sources, refusals, leaks, personal data, safety, latency, and cost for every answer.
A built-in adversarial suite, extended with cases for your own tools, data, and known risks.
Tests run on every change to prompts, retrieval, or models, and block the merge on failure.
Start from real questions in logs, tickets, or a pilot, with personal data removed.
Accuracy, grounding, refusal, permissions, injection, and safety, with critical cases marked.
Unique strings in restricted documents and the system prompt. Any answer containing one is a leak.
Pass rates overall and by category, cost and latency limits, and maximum drop against the last release.
The pipeline scores the change, publishes the report, and blocks the merge if a gate fails.
Bad answers seen in production go into the set before they are fixed.
Vendor-neutral Python and configuration, Azure first, with tests included from the start.
Thresholds are set in the Ground step and recorded in the system’s Evaluation Card.
Any critical case failing, or any case erroring, also fails the run. The report lists every case that passed last time and fails now.
From logs, tickets, or a pilot, with your experts writing the reference answers.
Pass rates, cost, and latency limits set with the business owner.
Your app's function, its API, or the model deployment on its own.
On every change to prompts, retrieval, or models, with the set refreshed each quarter.
You keep the test set, the gates, the pipeline, and a report for every release.
We state the limits up front, and we recommend tools based on fit. We do not resell platforms.
Get the 10-page PDF to share with your team, or tell us the decision you want to improve and we will tell you whether the GenAI Evaluation Tests fits.