# AI evaluation dataset

Build a collection of cases that makes versions comparable and exposes failures before they reach live operations. The quality of your reference answers matters as much as the model.

Revision: 2026-09-24

## When to use it

Before selecting an AI configuration, changing instructions or sources, or extending a project to new documents and teams.

## Expected output

A case register with verified expected results, judgement criteria and a separate set for the final evaluation.

## 1. Cover real work

Define the task before collecting examples. Include cases where asking for help or declining to answer is the correct behaviour.

### Task, input and required output

Specify what the system receives, what it should produce and actions it must not perform.

________________

### Categories to represent

List frequency, languages, formats, ambiguous cases, missing information and high-consequence situations.

________________

### Case sources and selection

Record period, permission to use the data, planned de-identification and limits of sample representativeness.

________________

> Illustrative example, replace with your own: Illustrative example: a deadline extraction task includes documents with no date and conflicting dates. For an ambiguous document, “review required” can be the correct result.

## 2. Create verifiable references

Give each case a stable identifier. Preserve the reason for its expected answer so disagreements can be resolved without relying on memory.

### Case identifier and input

Link the approved file or record and its version; avoid sensitive information in case names.

________________

### Expected output and evidence

Describe correct fields, supporting source, acceptable tolerances and behaviour when an answer is unavailable.

________________

### Reviewer and disagreements

Record who validates the reference and who resolves differences between evaluators.

________________

> Illustrative example, replace with your own: Illustrative example: DOC-017 contains two dates, neither labelled as the deadline. The reference requires no invented deadline and a pointer to the lines needing review.

## 3. Compare versions

Keep some cases apart from routine improvement work. Report results by category instead of hiding critical failures inside an overall average.

### Development and final evaluation

Record which cases may guide improvements and which remain separate for evaluation.

________________

### Measures and acceptance conditions

Define critical failures, field correctness, review requests and agreed limits on residual human effort.

________________

### Version, results and decision

Save configuration, date, category-level results, unresolved defects and the reason to release or defer.

________________

> Illustrative example, replace with your own: Illustrative example: a new configuration extracts more dates but invents a deadline in the ambiguous case. That critical defect remains visible even if the average improves.

## Pitfalls to avoid

- Using an unchecked AI answer as the reference result.

- Improving against the same small set and calling it an independent evaluation.

- Counting a well-formatted but unsupported answer as a success.

Stolen Orbit — https://stolenorbit.com/en/templates/ai-evaluation-dataset/

