Source: [Original HTML page](https://stolenorbit.com/en/resources/data-readiness-for-ai/)

Language: English

A PRACTICAL BUSINESS GUIDE

# Prepare data for AI by checking one process, not the entire archive.

You do not need to clean every company file to test feasibility. Start with the selected process, identify authoritative sources and examine the exceptions that could change its result.

[By Stolen Orbit](https://stolenorbit.com/en/about/) Updated 24 September 2026

## Define the data needed for one decision

[Source for this section](https://stolenorbit.com/en/resources/data-readiness-for-ai/#scope)

“We have lots of data” does not establish whether it can support the intended task. Preparing an order draft might require customer, product, quantity and unit. Confirming that order may additionally require commercial terms, availability and authorisation. These are different scopes with different requirements.

Write down the expected output and work backwards to the information that supports it. For each item, identify its location, owner and the consequence of its absence. This inventory distinguishes a blocking gap from an optional enhancement. Do not ask the model to guess information the business process needs to verify.

## Five checks before the first integration

[Source for this section](https://stolenorbit.com/en/resources/data-readiness-for-ai/#checks)

Scroll the table to compare all columns.

A source readiness record to complete for each source

| Check | Evidence needed | Example problem |
| --- | --- | --- |
| Access | Authorised account and permitted operations | The demonstration works only with an administrator account |
| Meaning | Field definitions and units | “Quantity 10” could mean items or packs |
| Quality | Complete, incomplete and conflicting examples | The same customer has different identifiers |
| Freshness | Owner and update frequency | An archived price list still appears current |
| Traceability | Source reference and version | Nobody can reconstruct where an amount came from |

## Build a sample that reveals problems

[Source for this section](https://stolenorbit.com/en/resources/data-readiness-for-ai/#sample)

Select recent cases and difficult cases rather than only the cleanest documents. Include different formats, rotated attachments, missing lines, abbreviations and new customers. An initial sample helps discover defects; it does not by itself establish future reliability across all production volumes.

Keep the correct outcome and its reasoning for every case. If two experienced people disagree, clarify the business rule first. Calling an undefined organisational decision an AI error makes evaluation unhelpful. Limit shared data to what the test requires and use an environment with suitable access controls. Record which formats or business units the sample does not cover.

## Illustrative example: ambiguous order quantities

[Source for this section](https://stolenorbit.com/en/resources/data-readiness-for-ai/#example)

1.  ### Observation
    
    A request says “3 boxes”, while the catalogue holds a price per individual item. Master data does not consistently record how many items a box contains.
    
2.  ### Decision
    
    The project can extract the wording and prepare a draft. It cannot automatically convert the quantity when an approved mapping is absent.
    
3.  ### Targeted repair
    
    The catalogue owner completes conversions for products included in the pilot. Other cases go to review with an explicit reason.
    

Data preparation is more than file cleaning. It includes meaning, ownership and behaviour when information is unavailable.

## Finish with practical lists and named owners

[Source for this section](https://stolenorbit.com/en/resources/data-readiness-for-ai/#outcome)

Avoid searching for a universal “AI readiness” score. The useful output is knowing which activities can be tested now and which depend on a specific repair. NIST stresses understanding context and documenting limitations; this worksheet is our operational application, not a certification. Review it again when a new data source or business unit is added.

-   Ready: usable sources, defined fields and tested access for the initial scope.
-   To repair: issues with an owner, concrete action and verification criterion.
-   Excluded: cases the system must recognise and send to a person.
-   To monitor: quality and freshness that may deteriorate after release.

FROM IDEAS TO A BRIEF

## A template to work from.

[Source for this section](https://stolenorbit.com/en/resources/data-readiness-for-ai/#downloads)

### AI data readiness worksheet

A source inventory with observed defects, remediation actions and an approved sample for the trial.

[Download the Markdown template](https://stolenorbit.com/downloads/data-readiness-en.md)

### Document extraction schema

A schema for each document family with field meanings, validation rules, evidence and missing-data behaviour.

[Download the Markdown template](https://stolenorbit.com/downloads/document-field-schema-en.md)

### Knowledge source register

An approved source catalogue with authoritative versions, permitted audiences, update cycles and retrieval tests.

[Download the Markdown template](https://stolenorbit.com/downloads/knowledge-source-register-en.md)

## Practical questions

[Source for this section](https://stolenorbit.com/en/resources/data-readiness-for-ai/#faq)

**Must we digitise everything first?**

No. Start with a scope that has sufficient sources. Document excluded inputs and do not automatically extend pilot results to different archives.

**Does the data have to be perfect?**

No, but errors must be recognisable and manageable. A missing field may trigger review; ambiguous information treated as certain can cause an incorrect decision.

**Who should take part in the assessment?**

Someone who performs the process, the source owner and whoever manages access or integrations. Include other specialists when the data and its intended use require them.

## References and method

[Source for this section](https://stolenorbit.com/en/resources/data-readiness-for-ai/#sources)

[NIST — AI Risk Management Framework 1.0: Core](https://airc.nist.gov/airmf-resources/airmf/5-sec-core/)

General reference on use context, measurement and documenting limitations. The checklist and example are Stolen Orbit operational proposals.

EXPLORE FURTHER

[AI data readiness worksheet](https://stolenorbit.com/en/templates/ai-data-readiness/)

[Document extraction schema](https://stolenorbit.com/en/templates/document-extraction-schema/)

[Knowledge source register](https://stolenorbit.com/en/templates/knowledge-source-register/)

[How to introduce AI](https://stolenorbit.com/en/resources/how-to-introduce-ai-in-your-business/)

[Turn documents into usable data, with checks at every handoff.](https://stolenorbit.com/en/services/ai-document-data-extraction/)

[Find answers in your company’s approved knowledge.](https://stolenorbit.com/en/services/internal-ai-knowledge-assistant/)

[All practical guides](https://stolenorbit.com/en/resources/)

## Which process should improve first?

[Source for this section](https://stolenorbit.com/en/resources/data-readiness-for-ai/#growth-closing-heading)

Start with a concrete process, the systems you use and the people who will operate it every day.

[Let’s discuss your process](https://stolenorbit.com/en/contact/)

## Turn the question into a next step.

-   [Use what you have, connect it, or build the missing part?](https://stolenorbit.com/en/resources/native-automation-connectors-custom-integrations/)
-   [Automation failures and duplicates](https://stolenorbit.com/en/resources/automation-failures-and-duplicates/)
-   [System integration map](https://stolenorbit.com/en/templates/integration-map/)
