Ninety-something is not an operating plan
A supplier says an AI tool is highly accurate. The team tests a handful of cases and likes the output. The pilot moves forward.
What remains unclear is whether the mistakes are harmless formatting issues, missing commitments or invented client facts. Those are different failures with different consequences. A single average can conceal the distinction even when it has been calculated honestly.
Our recommendation is to evaluate the workflow through a small error taxonomy before treating a headline score as a buying signal. Decide what can go wrong, how it is detected and what recovery requires.
Define the unit of success
For a document-extraction tool, success might mean a particular field is correct and tied to supporting evidence. For a meeting assistant, it might mean an agreed action is captured without inventing an owner. For search, it might mean the answer cites the relevant current document.
Do not combine these into a vague test of whether the output looks good. Build labelled examples of the actual task, including cases where the correct response is uncertainty or refusal to infer.
The ICO distinguishes personal-data accuracy from statistical accuracy and cautions that overall model accuracy may not be a useful enough measure. Our proposed error categories apply that distinction to operational testing; they are not an ICO-mandated scorecard.
Cost the failures separately
Classify errors by the work or harm they could create. A missed action may delay a case. An invented action may send staff on unnecessary work. An incorrect client fact may affect a later decision if it is accepted into the record.
For each class, identify the detection point, reviewer and recovery route. Estimate review effort during planning, then replace estimates with observations from the pilot. Do not present projected savings as measured results.
Record near misses too. If a reviewer catches an error, that demonstrates a control working. It does not make the underlying model output correct.
Test beyond the demonstration
Include incomplete documents, revised information, ambiguous wording and unusual but legitimate cases. Keep a separate set of examples for final evaluation so repeated tuning does not turn the test into a memorised examination.
Evaluate the whole process: retrieval, model output, review and any resulting action. A model may perform well on supplied text while the live system selects the wrong evidence. The client-facing result depends on both.
NIST's AI Risk Management Framework is voluntary and provides a general framework for AI risk work. It does not certify a firm's pilot or supply a universal passing score. Use such frameworks to organise thinking, then set task-specific acceptance conditions.
Make release conditional
Agree which failures block release, which require human review and which can be corrected through ordinary support. Decide this before the team becomes attached to the pilot's apparent success.
After a model, prompt or retrieval change, rerun the relevant examples. Keep a record of what changed and compare error types, not just the overall result. Improvements in one category can coexist with deterioration in another.
The number that belongs in the business case
The final question is whether the firm can detect and absorb the errors while still gaining useful capacity. A tool that saves drafting time but consumes more review time has not yet made the workflow better.
Sources & further reading
- AI accuracy and statistical accuracy · accessed 2026-09-13
- NIST AI Risk Management Framework · accessed 2026-09-13
Recommendations and examples are editorial analysis, not personalised financial or legal advice. Source links allow readers to check the underlying evidence.