# AI operations runbook

Make it clear who watches the system and what they do when work stops or quality declines. Monitoring includes detecting missing cases, not only reported errors.

Revision: 2026-09-24

## When to use it

Before production handover and when defining ongoing support between your business and its implementation partner.

## Expected output

Instructions available to the responsible team: signals, alerts, initial diagnosis, recovery and periodic review.

## 1. Decide what to observe

Connect signals to an operational outcome. A process can respond successfully without completing work or receiving all expected inputs.

### Workflow health and expected inputs

Record last event, cases received and completed, queue size, latency and checks for unexpectedly absent inputs.

________________

### Quality and residual work

Define checks for errors, human corrections, missing sources and cases routed to manual handling.

________________

### Usage and operational limits

Specify observable volume and costs, who investigates changes and which limits require action.

________________

> Illustrative example, replace with your own: Illustrative example: no errors are logged, but zero requests are processed during a normally active morning. Check the inbox connection before concluding that the service is healthy.

## 2. Make alerts actionable

Every alert needs a recipient and a first action. Define agreed coverage and response windows without assuming the team is always available.

### Signal, threshold and recipient

For each alert record an observable condition, the role notified and a monitored channel.

________________

### Initial checks and access

List dashboards, log references and non-destructive checks; do not include secrets.

________________

### Escalation and containment

Specify when to contact the next owner, restrict the workflow or start manual handling.

________________

> Illustrative example, replace with your own: Illustrative example: when a queue grows, check the ERP connection and reviewer capacity. If the ERP is blocked, pause writes while preserving queued work.

## 3. Maintain the service

A backup operator must be able to use the runbook. Periodically confirm that links, permissions, sources and named people remain valid.

### Recovery and final checks

Explain how to resume, reconcile pending cases and inspect partial effects or duplicates.

________________

### Reviews and planned maintenance

List checks for versions, access, sources, evaluation cases and service capacity.

________________

### Intervention log and updates

Record where incidents and changes are documented and who maintains these instructions.

________________

> Illustrative example, replace with your own: Illustrative example: after restoring ERP access, compare requests with created drafts before draining the queue. Add the successful recovery step to the runbook.

## Pitfalls to avoid

- Sending alerts to an inbox nobody monitors.

- Measuring technical availability without checking completion and quality.

- Agreeing ongoing support without defining coverage, boundaries and ownership.

Stolen Orbit — https://stolenorbit.com/en/templates/ai-monitoring-runbook/

