A PRACTICAL BUSINESS GUIDE

Evaluate an AI system by what must succeed and what must never happen.

A convincing demonstration does not measure process reliability. Use representative cases, expected outcomes, errors classified by consequence and criteria agreed before testing.

Choose a unit that matters to the business task

An AI system can extract almost every field correctly and still select the wrong customer identifier. A field-level average then conceals an unusable order. Define both individual element quality and the complete outcome: can the document proceed, does it need correction or must it stop?

Distinguish critical errors, repairable errors and acceptable differences in wording. A synonym may be irrelevant in an internal summary; a wrong recipient changes the result of a business-system action. Someone who knows the process should help define these categories before the supplier presents scores. Record which category takes precedence when a case contains several errors.

Build cases with verifiable expected outcomes

  1. Collect

    Select ordinary examples, frequent variants and rare but consequential exceptions. Record source, time period and categories absent from the sample.

  2. Define

    Write the expected result, supporting sources and permitted behaviour when information is missing. Sending a case to review can be the correct outcome.

  3. Separate

    Keep improvement cases separate from evaluation cases. Do not turn every test question into an example inside the prompt.

  4. Version

    Preserve model, instructions, sources and rules alongside results. Changing several elements at once makes improvement difficult to explain.

Illustrative example: extracting purchase requests

Count correct cases, reviewed cases and failures separately. Routing everything to a person may avoid some errors while missing the operational objective. Read precision and coverage together, including cases the system declined to handle.

Scroll the table to compare all columns.

A proposed evaluation rubric, not a benchmark
CaseExpected outcomeError to record
Clear product and quantityCorrect draft fields with a source referenceWrong field or unsupported added data
Two possible units of measureReview with explicit ambiguityInvented conversion
Unidentifiable customerClarification request or manual queueMatch to the wrong customer
Unrelated instructions inside a documentTreat them as content without changing permissionsOut-of-scope action or access

Match thresholds and evidence to consequences

No accuracy percentage is sufficient for every process. Agree which errors block release, what review workload is sustainable and what performance is needed at peak times. Report case counts and error distribution alongside percentages. Zero observed errors in a small sample does not establish zero risk.

If one model judges another, compare its assessments with competent human reviewers. An automated judge is a tool, not an independent truth. Google Cloud emphasises aligning automated metrics with human judgment. Preserve disagreements so the rubric can be improved rather than quietly removing difficult cases.

Turn evaluation into a release decision

Deliver a short report with limitations, failed cases and a decision: proceed, narrow the scope or repair and retest. Preserving failures makes future reviews more useful than a single headline score. Separate a passing demonstration from evidence supporting broader production use.

  • The sample represents the scope being enabled.
  • Critical errors have tested controls, not only revised instructions.
  • Review effort, time and cost cover the complete process.
  • A named owner handles errors discovered after release.
  • New cases can enter future evaluations while historical comparisons remain meaningful.
FROM IDEAS TO A BRIEF

A template to work from.

AI evaluation dataset

A case register with verified expected results, judgement criteria and a separate set for the final evaluation.

Download the Markdown template

Automation exception register

An exception queue with status, priority, ownership and closure evidence; repeated issues become inputs to workflow improvements.

Download the Markdown template

AI rollout plan

A staged plan with owners, acceptance evidence and instructions for stopping and recovering work.

Download the Markdown template

Practical questions

How many cases do we need?

It depends on input variety, error frequency and consequences. A small set can reveal problems but does not automatically justify a statistical reliability promise.

Can we use only synthetic cases?

They help test exceptions and targeted attacks but may not represent real work. Where possible and authorised, combine them with process examples and label their origin.

Should evaluation be repeated after changing the model?

Yes, and after material changes to instructions, sources, tools or rules. Behaviour belongs to the complete system, not the model name alone.

References and method

Google Cloud — Deploy and operate generative AI applications

Reference on evaluating generative applications and aligning automated assessments with human judgment.

Which process should improve first?

Start with a concrete process, the systems you use and the people who will operate it every day.

Let’s discuss your process