Available knowledge and behaviour are different requirements
Retrieval augmented generation, or RAG, searches external sources and supplies information to the model while it answers. Fine tuning adapts a model using training examples. Microsoft distinguishes retrieving knowledge from adapting behaviour, style or task performance. The techniques can be combined.
This does not mean every chatbot needs a vector database or dedicated training. A small set of instructions and documents may be enough for an initial test. Before adding components, keep a record of questions, expected answers and failure reasons: missing source, incorrect retrieval, wrong interpretation or failure to follow the required format.
Match the intervention to an observed defect
Scroll the table to compare all columns.
| Defect | Initial hypothesis | What to check |
|---|---|---|
| A recent procedure is missing | Retrieval from approved sources | Date, version, permissions and the passage actually retrieved |
| Answers are accurate but formatting varies | Instructions, examples and validation; consider fine tuning if needed | Schema compliance on cases never used for optimisation |
| Stock quantity or order status is wrong | Read the authoritative operational system | Identifier, freshness and access rules |
| Documents contradict one another | Resolve and govern the sources | Which document takes precedence and who approves it |
Illustrative example: an internal procedure assistant
Suppose an operations team asks which documents are needed to open a case. The procedure changes while older versions remain in shared folders. Training a model on available files does not establish who owns the current version. Assign a procedure owner and make the valid version identifiable first.
In the test, the assistant should show the passage supporting its answer and recognise when the source lacks the information. An existing citation is insufficient: the reviewer must check whether it supports the particular claim. A question about a specific customer may instead require an authorised read from the business system, rather than a search through the procedure.
Evaluate retrieval separately from the answer
When comparing configurations, change one element at a time and preserve versions. Improvement may come from cleaner documents rather than a newer model. This distinction tells you where to invest the next unit of effort. A retrieval defect and an answer-generation defect require different fixes.
- Record the document and passage that should be found for each question.
- Check whether retrieval returns the right source before judging the generated wording.
- Test unanswerable questions, outdated documents and users with different permissions.
- Keep evaluation examples separate from those used to change prompts or training.
- Measure source maintenance, waiting time and review work as well as answer quality.
What to request from the project
Request a source register, a set of evaluation questions and an explanation of excluded material. Ask how deletions or revoked permissions reach the search system and how long propagation takes. For fine tuning, establish who maintains examples and how a new version is checked.
A valid evaluation outcome can be “fine tuning is unnecessary” or “fix retrieval before changing the model”. The selected technology should address a measured error. “Trained on your data” is not a sufficient description of sources, access or updates. Ask to inspect one complete example from source selection through the final answer before approving a broader rollout.
A template to work from.
Knowledge source register
An approved source catalogue with authoritative versions, permitted audiences, update cycles and retrieval tests.
Download the Markdown templateAI evaluation dataset
A case register with verified expected results, judgement criteria and a separate set for the final evaluation.
Download the Markdown templateAI data readiness worksheet
A source inventory with observed defects, remediation actions and an approved sample for the trial.
Download the Markdown templatePractical questions
Does RAG eliminate invented answers?
No. Retrieval can miss the source, find the wrong document or supply information the model misinterprets. Testing and an explicit path for insufficient evidence remain necessary.
Is fine tuning necessary to use company documents?
Not necessarily. Documents and data can be supplied in context or retrieved from external sources. Test the choice against the task and its observed errors.
What about data that changes every minute?
For current operational status, consider a controlled read from the authoritative system. Define freshness and behaviour when that system is unavailable instead of treating an old copy as current data.
References and method
Definitions of retrieval, indexes and model adaptation. The scenarios and test criteria are original operational examples.