Source: [Original HTML page](https://stolenorbit.com/en/resources/evaluating-ai-reliability/)

Language: English

A PRACTICAL BUSINESS GUIDE

# Evaluate an AI system by what must succeed and what must never happen.

A convincing demonstration does not measure process reliability. Use representative cases, expected outcomes, errors classified by consequence and criteria agreed before testing.

[By Stolen Orbit](https://stolenorbit.com/en/about/) Updated 24 September 2026

## Choose a unit that matters to the business task

[Source for this section](https://stolenorbit.com/en/resources/evaluating-ai-reliability/#unit)

An AI system can extract almost every field correctly and still select the wrong customer identifier. A field-level average then conceals an unusable order. Define both individual element quality and the complete outcome: can the document proceed, does it need correction or must it stop?

Distinguish critical errors, repairable errors and acceptable differences in wording. A synonym may be irrelevant in an internal summary; a wrong recipient changes the result of a business-system action. Someone who knows the process should help define these categories before the supplier presents scores. Record which category takes precedence when a case contains several errors.

## Build cases with verifiable expected outcomes

[Source for this section](https://stolenorbit.com/en/resources/evaluating-ai-reliability/#dataset)

1.  ### Collect
    
    Select ordinary examples, frequent variants and rare but consequential exceptions. Record source, time period and categories absent from the sample.
    
2.  ### Define
    
    Write the expected result, supporting sources and permitted behaviour when information is missing. Sending a case to review can be the correct outcome.
    
3.  ### Separate
    
    Keep improvement cases separate from evaluation cases. Do not turn every test question into an example inside the prompt.
    
4.  ### Version
    
    Preserve model, instructions, sources and rules alongside results. Changing several elements at once makes improvement difficult to explain.
    

## Illustrative example: extracting purchase requests

[Source for this section](https://stolenorbit.com/en/resources/evaluating-ai-reliability/#example)

Count correct cases, reviewed cases and failures separately. Routing everything to a person may avoid some errors while missing the operational objective. Read precision and coverage together, including cases the system declined to handle.

Scroll the table to compare all columns.

A proposed evaluation rubric, not a benchmark

| Case | Expected outcome | Error to record |
| --- | --- | --- |
| Clear product and quantity | Correct draft fields with a source reference | Wrong field or unsupported added data |
| Two possible units of measure | Review with explicit ambiguity | Invented conversion |
| Unidentifiable customer | Clarification request or manual queue | Match to the wrong customer |
| Unrelated instructions inside a document | Treat them as content without changing permissions | Out-of-scope action or access |

## Match thresholds and evidence to consequences

[Source for this section](https://stolenorbit.com/en/resources/evaluating-ai-reliability/#thresholds)

No accuracy percentage is sufficient for every process. Agree which errors block release, what review workload is sustainable and what performance is needed at peak times. Report case counts and error distribution alongside percentages. Zero observed errors in a small sample does not establish zero risk.

If one model judges another, compare its assessments with competent human reviewers. An automated judge is a tool, not an independent truth. Google Cloud emphasises aligning automated metrics with human judgment. Preserve disagreements so the rubric can be improved rather than quietly removing difficult cases.

## Turn evaluation into a release decision

[Source for this section](https://stolenorbit.com/en/resources/evaluating-ai-reliability/#release)

Deliver a short report with limitations, failed cases and a decision: proceed, narrow the scope or repair and retest. Preserving failures makes future reviews more useful than a single headline score. Separate a passing demonstration from evidence supporting broader production use.

-   The sample represents the scope being enabled.
-   Critical errors have tested controls, not only revised instructions.
-   Review effort, time and cost cover the complete process.
-   A named owner handles errors discovered after release.
-   New cases can enter future evaluations while historical comparisons remain meaningful.

FROM IDEAS TO A BRIEF

## A template to work from.

[Source for this section](https://stolenorbit.com/en/resources/evaluating-ai-reliability/#downloads)

### AI evaluation dataset

A case register with verified expected results, judgement criteria and a separate set for the final evaluation.

[Download the Markdown template](https://stolenorbit.com/downloads/evaluation-dataset-en.md)

### Automation exception register

An exception queue with status, priority, ownership and closure evidence; repeated issues become inputs to workflow improvements.

[Download the Markdown template](https://stolenorbit.com/downloads/exception-register-en.md)

### AI rollout plan

A staged plan with owners, acceptance evidence and instructions for stopping and recovering work.

[Download the Markdown template](https://stolenorbit.com/downloads/rollout-plan-en.md)

## Practical questions

[Source for this section](https://stolenorbit.com/en/resources/evaluating-ai-reliability/#faq)

**How many cases do we need?**

It depends on input variety, error frequency and consequences. A small set can reveal problems but does not automatically justify a statistical reliability promise.

**Can we use only synthetic cases?**

They help test exceptions and targeted attacks but may not represent real work. Where possible and authorised, combine them with process examples and label their origin.

**Should evaluation be repeated after changing the model?**

Yes, and after material changes to instructions, sources, tools or rules. Behaviour belongs to the complete system, not the model name alone.

## References and method

[Source for this section](https://stolenorbit.com/en/resources/evaluating-ai-reliability/#sources)

[Google Cloud — Deploy and operate generative AI applications](https://docs.cloud.google.com/architecture/deploy-operate-generative-ai-applications)

Reference on evaluating generative applications and aligning automated assessments with human judgment.

EXPLORE FURTHER

[AI evaluation dataset](https://stolenorbit.com/en/templates/ai-evaluation-dataset/)

[Automation exception register](https://stolenorbit.com/en/templates/exception-register/)

[AI rollout plan](https://stolenorbit.com/en/templates/ai-rollout-plan/)

[AI agents for business](https://stolenorbit.com/en/services/ai-agents/)

[Turn documents into usable data, with checks at every handoff.](https://stolenorbit.com/en/services/ai-document-data-extraction/)

[Find answers in your company’s approved knowledge.](https://stolenorbit.com/en/services/internal-ai-knowledge-assistant/)

[All practical guides](https://stolenorbit.com/en/resources/)

## Which process should improve first?

[Source for this section](https://stolenorbit.com/en/resources/evaluating-ai-reliability/#growth-closing-heading)

Start with a concrete process, the systems you use and the people who will operate it every day.

[Let’s discuss your process](https://stolenorbit.com/en/contact/)

## Turn the question into a next step.

-   [Evaluation lab](https://stolenorbit.com/en/tools/ai-extraction-evaluation/)
-   [AI evaluation dataset](https://stolenorbit.com/en/templates/ai-evaluation-dataset/)
-   [Human review for AI](https://stolenorbit.com/en/resources/human-review-for-ai/)
