← The archive
OPINION

Your AI pilot needs a boring success measure.

A polished demo is a poor substitute for a useful result.

· 2 min read

Opinion · An editorial recommendation

Demos measure the demo

A pilot that ends with an impressive session proves the tool can produce something plausible under favourable conditions. That is not the claim anyone is buying, which is that a real piece of work gets better.

So define the piece of work first, then the measure, then run the tool against it. In that order the result is interpretable either way.

Include the checking in the cost

The time spent verifying and correcting output is part of the task, not overhead outside it. A draft produced in seconds that needs a specialist twenty minutes to make safe has moved effort rather than removed it — sometimes usefully, sometimes not.

Measure the whole path from request to something you would put your name on. Compare it against the same task done the existing way by someone of similar experience.

Decide what data the pilot may use

Agree in advance which material is approved for the tool and where the output may be stored. This is easier to settle before people have built a habit around it, and it protects the pilot from being cancelled for reasons unrelated to its results.

Stopping is a legitimate outcome

A pilot that ends with a clear “not for this task, and here is the evidence” has done its job. The failure mode is the pilot that never concludes and quietly becomes a subscription.

Write the decision down with the numbers behind it. It saves the same conversation happening again in six months.

An example with the checking included

Imagine a fictional team testing an assistant on internal meeting summaries. Choose a small, varied set of approved source material: a clear meeting, one with conflicting views and one with several decisions. Agree what the summary must preserve, including actions, owners and unresolved questions.

For each case, record the time to produce and verify the summary, missed or invented actions and the amount of rewriting required. These are example measures, not a research result. They let you see whether a quick first draft leads to a useful final document.

Use a decision rule before seeing the result

Agree the conditions for extending the pilot, changing its scope or stopping. A team might require less total handling time without an increase in material errors, with every output still reviewed by a named person. Choose the threshold around the consequences of the task rather than borrowing a convenient percentage.

Keep a short decision record: task tested, approved data, tool configuration, baseline, results, limitations and next owner. If the sample was small or unusually tidy, say so. A promising result can justify another bounded test without justifying an organisation-wide purchase.

Something needs correcting?

Get in touch to explain which claim needs attention and share the supporting source.