Use the same boundary before and after
Define when a case starts and when it is complete. Comparing data-entry time before with data entry plus review afterwards is inconsistent. Excluding corrections afterwards makes the improvement look larger than it is. Use the same boundary and separate active work from elapsed waiting time.
Collect cases under representative conditions, noting volume, category, operator and complexity where relevant. Do not compare a quiet period with a peak without disclosing it. If you cannot isolate automation’s effect, describe the comparison as an observation and identify other changes that happened. Preserve the original baseline rather than replacing it each time results improve.
A small but complete operating dashboard
Scroll the table to compare all columns.
| Metric | How to read it | What it can hide alone |
|---|---|---|
| Coverage | Handled cases / eligible cases in the period | Exclusion of difficult cases |
| Quality | Cases accepted without correction / handled cases | Rare but serious errors |
| Active work | Preparation, review and recovery minutes | Waiting and customer interruptions |
| Lead time | Arrival through actual completion | A bottleneck moved elsewhere |
| Backlog | Open cases, age and priority | A small queue containing very old cases |
| Business outcome | The specific result agreed before testing | Changes caused by other factors |
Illustrative example: released capacity, but a queue still waits
Assume arbitrary inputs: 100 cases previously took 10 minutes of work each, totalling 1,000 minutes. After the change, 70 cases need four minutes of review, 30 remain manual at ten minutes and error recovery adds 120 minutes. Total work is now 700 minutes: 300 minutes, or five hours, of released capacity.
This does not establish five hours of eliminated cost. The team may use them for backlog, quality or additional volume. If every case still waits two days for a signature, customer-perceived lead time may barely change. The next decision may concern that approval step rather than making the model faster.
This is an educational example, not a client result. Categories must be distinct to avoid counting corrections or manual work twice.
Compare segments and distributions, not only averages
Separate formats, case types or teams working under different conditions. Report median and slowest-case behaviour when helpful for understanding delays. A better average can coexist with worse handling of the most important exceptions.
Maintain a change log: revised prompt, updated source, changed procedure or trained team member. Connect changes to observations without automatically claiming causation. Where feasible, compare similar groups or introduce a change in stages while documenting differences and limitations. Record missing data as missing rather than silently excluding inconvenient cases from the denominator.
Use measurements to choose the next action
NIST includes measurement and management throughout the AI lifecycle. Our practical proposal is an operational review with an owner, readable evidence and a recorded decision. Use the separate cost guide and worksheet to translate benefits and expenses into economic scenarios. Keep the operational record available so those scenarios can later be compared with actual outcomes.
- If coverage is low, identify excluded inputs and their reasons.
- If corrections are frequent, separate data, instruction and process problems.
- If effort falls but delays remain, find the next bottleneck.
- If the review backlog grows, narrow the scope or increase review capacity.
- If capacity is released, explicitly assign its new use and measure that outcome.
A template to work from.
Business reporting brief
A report contract covering audience, metric definitions, sources, checks, reviewer and distribution conditions.
Download the Markdown templateAI change log
A version history with rationale, impact, verification evidence and a release decision.
Download the Markdown templateAutomation opportunity scorecard
A reasoned comparison, including missing evidence, blocking conditions and a decision about the next experiment.
Download the Markdown templatePractical questions
Are hours saved an ROI?
They are a possible operational benefit. An economic return also requires complete costs and an explanation of how capacity becomes realised value. It does not automatically mean reduced payroll or spending.
How many metrics should we track?
Enough to support a concrete decision. Start with coverage, quality, complete effort, waiting time and business outcome; add detail when it explains a problem or guides an action.
How do we measure a drafting assistant?
Record time to a usable version, corrections, material errors and actual use of the draft. The number of generated texts alone does not measure the outcome.
References and method
General reference on lifecycle measurement and management. Metrics, formulas and examples here are original operational proposals.