FREE TEMPLATE · FILL IN, DOWNLOAD, REUSE

AI evaluation dataset

Build a collection of cases that makes versions comparable and exposes failures before they reach live operations. The quality of your reference answers matters as much as the model.

When to use it and what to produce

Before selecting an AI configuration, changing instructions or sources, or extending a project to new documents and teams.

A case register with verified expected results, judgement criteria and a separate set for the final evaluation.

Fill in the fields below and download your document. The website does not send or save answers: closing the page may lose them. Examples are illustrative and should be replaced with your process data.

Download the complete blank template (.md)

1. Cover real work

Define the task before collecting examples. Include cases where asking for help or declining to answer is the correct behaviour.

See a worked example

Illustrative example: a deadline extraction task includes documents with no date and conflicting dates. For an ambiguous document, “review required” can be the correct result.

Specify what the system receives, what it should produce and actions it must not perform.

List frequency, languages, formats, ambiguous cases, missing information and high-consequence situations.

Record period, permission to use the data, planned de-identification and limits of sample representativeness.

2. Create verifiable references

Give each case a stable identifier. Preserve the reason for its expected answer so disagreements can be resolved without relying on memory.

See a worked example

Illustrative example: DOC-017 contains two dates, neither labelled as the deadline. The reference requires no invented deadline and a pointer to the lines needing review.

Link the approved file or record and its version; avoid sensitive information in case names.

Describe correct fields, supporting source, acceptable tolerances and behaviour when an answer is unavailable.

Record who validates the reference and who resolves differences between evaluators.

3. Compare versions

Keep some cases apart from routine improvement work. Report results by category instead of hiding critical failures inside an overall average.

See a worked example

Illustrative example: a new configuration extracts more dates but invents a deadline in the ambiguous case. That critical defect remains visible even if the average improves.

Record which cases may guide improvements and which remain separate for evaluation.

Define critical failures, field correctness, review requests and agreed limits on residual human effort.

Save configuration, date, category-level results, unresolved defects and the reason to release or defer.

Pitfalls to avoid

  • Using an unchecked AI answer as the reference result.
  • Improving against the same small set and calling it an independent evaluation.
  • Counting a well-formatted but unsupported answer as a success.

Take the next step together.

Use this document to start a discussion about scope, evaluation and project ownership.

Review the text before sending. Only your email is needed alongside the message; other details are optional.