FREE TEMPLATE · FILL IN, DOWNLOAD, REUSE

AI operations runbook

Make it clear who watches the system and what they do when work stops or quality declines. Monitoring includes detecting missing cases, not only reported errors.

When to use it and what to produce

Before production handover and when defining ongoing support between your business and its implementation partner.

Instructions available to the responsible team: signals, alerts, initial diagnosis, recovery and periodic review.

Fill in the fields below and download your document. The website does not send or save answers: closing the page may lose them. Examples are illustrative and should be replaced with your process data.

Download the complete blank template (.md)

1. Decide what to observe

Connect signals to an operational outcome. A process can respond successfully without completing work or receiving all expected inputs.

See a worked example

Illustrative example: no errors are logged, but zero requests are processed during a normally active morning. Check the inbox connection before concluding that the service is healthy.

Record last event, cases received and completed, queue size, latency and checks for unexpectedly absent inputs.

Define checks for errors, human corrections, missing sources and cases routed to manual handling.

Specify observable volume and costs, who investigates changes and which limits require action.

2. Make alerts actionable

Every alert needs a recipient and a first action. Define agreed coverage and response windows without assuming the team is always available.

See a worked example

Illustrative example: when a queue grows, check the ERP connection and reviewer capacity. If the ERP is blocked, pause writes while preserving queued work.

For each alert record an observable condition, the role notified and a monitored channel.

List dashboards, log references and non-destructive checks; do not include secrets.

Specify when to contact the next owner, restrict the workflow or start manual handling.

3. Maintain the service

A backup operator must be able to use the runbook. Periodically confirm that links, permissions, sources and named people remain valid.

See a worked example

Illustrative example: after restoring ERP access, compare requests with created drafts before draining the queue. Add the successful recovery step to the runbook.

Explain how to resume, reconcile pending cases and inspect partial effects or duplicates.

List checks for versions, access, sources, evaluation cases and service capacity.

Record where incidents and changes are documented and who maintains these instructions.

Pitfalls to avoid

  • Sending alerts to an inbox nobody monitors.
  • Measuring technical availability without checking completion and quality.
  • Agreeing ongoing support without defining coverage, boundaries and ownership.

Take the next step together.

Use this document to start a discussion about scope, evaluation and project ownership.

Review the text before sending. Only your email is needed alongside the message; other details are optional.