Choose a unit that matters to the business task
An AI system can extract almost every field correctly and still select the wrong customer identifier. A field-level average then conceals an unusable order. Define both individual element quality and the complete outcome: can the document proceed, does it need correction or must it stop?
Distinguish critical errors, repairable errors and acceptable differences in wording. A synonym may be irrelevant in an internal summary; a wrong recipient changes the result of a business-system action. Someone who knows the process should help define these categories before the supplier presents scores. Record which category takes precedence when a case contains several errors.
Build cases with verifiable expected outcomes
Collect
Select ordinary examples, frequent variants and rare but consequential exceptions. Record source, time period and categories absent from the sample.
Define
Write the expected result, supporting sources and permitted behaviour when information is missing. Sending a case to review can be the correct outcome.
Separate
Keep improvement cases separate from evaluation cases. Do not turn every test question into an example inside the prompt.
Version
Preserve model, instructions, sources and rules alongside results. Changing several elements at once makes improvement difficult to explain.
Illustrative example: extracting purchase requests
Count correct cases, reviewed cases and failures separately. Routing everything to a person may avoid some errors while missing the operational objective. Read precision and coverage together, including cases the system declined to handle.
Scroll the table to compare all columns.
| Case | Expected outcome | Error to record |
|---|---|---|
| Clear product and quantity | Correct draft fields with a source reference | Wrong field or unsupported added data |
| Two possible units of measure | Review with explicit ambiguity | Invented conversion |
| Unidentifiable customer | Clarification request or manual queue | Match to the wrong customer |
| Unrelated instructions inside a document | Treat them as content without changing permissions | Out-of-scope action or access |
Match thresholds and evidence to consequences
No accuracy percentage is sufficient for every process. Agree which errors block release, what review workload is sustainable and what performance is needed at peak times. Report case counts and error distribution alongside percentages. Zero observed errors in a small sample does not establish zero risk.
If one model judges another, compare its assessments with competent human reviewers. An automated judge is a tool, not an independent truth. Google Cloud emphasises aligning automated metrics with human judgment. Preserve disagreements so the rubric can be improved rather than quietly removing difficult cases.
Turn evaluation into a release decision
Deliver a short report with limitations, failed cases and a decision: proceed, narrow the scope or repair and retest. Preserving failures makes future reviews more useful than a single headline score. Separate a passing demonstration from evidence supporting broader production use.
- The sample represents the scope being enabled.
- Critical errors have tested controls, not only revised instructions.
- Review effort, time and cost cover the complete process.
- A named owner handles errors discovered after release.
- New cases can enter future evaluations while historical comparisons remain meaningful.
A template to work from.
AI evaluation dataset
A case register with verified expected results, judgement criteria and a separate set for the final evaluation.
Download the Markdown templateAutomation exception register
An exception queue with status, priority, ownership and closure evidence; repeated issues become inputs to workflow improvements.
Download the Markdown templateAI rollout plan
A staged plan with owners, acceptance evidence and instructions for stopping and recovering work.
Download the Markdown templatePractical questions
How many cases do we need?
It depends on input variety, error frequency and consequences. A small set can reveal problems but does not automatically justify a statistical reliability promise.
Can we use only synthetic cases?
They help test exceptions and targeted attacks but may not represent real work. Where possible and authorised, combine them with process examples and label their origin.
Should evaluation be repeated after changing the model?
Yes, and after material changes to instructions, sources, tools or rules. Behaviour belongs to the complete system, not the model name alone.
References and method
Reference on evaluating generative applications and aligning automated assessments with human judgment.