How to measure whether an AI pilot actually worked

TopicMethod Published2026-08-23 Read8 min

A pilot that started without a baseline cannot prove anything, because there is nothing to subtract from. UK government benefits management doctrine, published in the Teal Book, requires that current performance is captured before work starts, that every benefit has a named owner, and that monetisable benefits are separated into cash-releasing and non-cash. Applied to an AI pilot, that means hours released and cash converted are different figures that must never be added together, and anything a vendor claims stays in a third column until a baseline makes it measurable.

A pilot finishes. Someone asks what it saved. The vendor says a hundred and nineteen thousand pounds. The team lead says it definitely feels faster. Finance says it cannot see anything in the numbers. All three are being honest, and there is no way to settle it, because nobody wrote down what things looked like before.

This happens constantly, and it is the reason so many AI pilots end in a shrug rather than a decision. The fix is not complicated and it is not new. The UK government wrote it down, it is free to read, and almost nobody in HR has read it.

The rule everything else hangs off

The Teal Book is the government’s project delivery guidance. Chapter 19 covers benefits management. On baselines it says this:

“The current performance level for each benefit should be established before work starts, and captured on the benefits register, providing the baseline for measuring benefits realisation.”

Before work starts. Not during, not retrospectively, not from memory.

That single sentence is the difference between a pilot that produces a decision and a pilot that produces an argument. If you cannot say what a task cost in time or money in the month before the tool arrived, you cannot say what the tool did. You can say people liked it. You cannot say it worked.

The same chapter defines a benefit owner as:

“a named individual who confirms the benefit’s identification and value, agrees the plan in the business case and then takes responsibility for realising and reporting on it.”

A named individual. Not a function, not a steering group. In practice this should almost always be the line manager whose numbers will move, because they are the only person who can actually make the benefit happen and the only person who will notice if it does not.

What a benefit profile has to contain

Government guidance expects each benefit to be documented with the measure, the owner, the baseline, the target, the timescale, the assumptions and dependencies, the risks, the expected path from baseline to target, and the realisation milestones.

That looks like a lot until you try to run a pilot without it. Every one of those fields exists because somebody once lost an argument for want of it.

The two that get skipped most often are the assumptions and the trajectory. The assumptions matter because a benefit case usually rests on one shaky belief that nobody wrote down, such as “volumes will stay flat” or “the team will not be restructured”. The trajectory matters because benefits rarely arrive in a straight line, and if you have not said when you expect the curve to bend, any interim reading looks like failure.

The distinction that ends most disagreements

Government guidance splits benefits into monetisable and unmonetisable. Then it splits monetisable again, into cash-releasing and non-cash.

Cash-releasing means it affects income or spending. Non-cash means it does not.

Apply that to an AI pilot and the fog clears immediately.

Say your People Operations team handles 400 first-line policy questions a month. Each one takes about twelve minutes end to end. You put a tool in front of it and the average drops to three minutes.

Nine minutes saved, four hundred times, is sixty hours a month. At a loaded cost of twenty two pounds an hour that is £1,320 a month, or £15,840 a year.

Here is where most business cases go wrong. That £15,840 is not a saving. It is sixty hours a month of released capacity, and released capacity only becomes money when somebody makes a decision about it. A vacancy goes unfilled. Agency cover stops. A project that would have needed a contractor gets absorbed. Until one of those happens, your finance director will look for the saving, fail to find it, and reasonably conclude the pilot did nothing.

So you report two numbers, not one.

Hours released: 720 a year. Measured, from ticket timestamps. Cash converted: nil so far. Because no budget decision has been taken yet.

Those are both true, they do not contradict each other, and neither of them is inflated. That is a report a finance director will trust, and trust is the thing you actually need if you want a second pilot funded.

The third column

There is one more figure worth tracking separately, and it is the one that keeps everybody honest.

Whatever the vendor says the tool will deliver. Whatever the business case claimed at approval. The number in the slide deck.

Keep it. Do not throw it away and do not fold it into the others. Label it as claimed and not yet realised, and let the gap between it and the measured figures be visible for the whole programme. Two things follow. Your forecasting gets better, because you can see by how much you tend to over-claim. And the conversation with the supplier changes, because the gap is on the table rather than in somebody’s head.

Three columns, kept apart, never summed:

What it meansWhere it comes from
Hours releasedNon-cash capacity freedTime study, system timestamps
Cash convertedBudget actually changedFinance ledger, signed off
Claimed, not realisedAsserted, not yet provenBusiness case or vendor

Adding them together produces a number that will not survive its first meeting with anybody numerate. Keeping them apart produces something you can defend in front of an audit committee.

What to do before the next pilot starts

Four things, and they take about half a day between them.

Write down the measure in the customer’s own words first, then work out how to count it. “Every candidate hears back within five working days” is a better starting point than “reduce time to response”, because it is checkable and because somebody actually cares about it.

Find where the number already lives. Applicant tracking timestamps, helpdesk tickets, payroll, case management. Most People functions are sitting on more usable baseline data than they think, and it is usually in a system nobody associates with measurement.

Name the owner and tell them. Not a group. A person, who knows they have been named, and who has the authority to change the thing being measured.

Then agree, before anything goes live, what result would cause you to stop. A pilot with no stopping condition is not a pilot. It is a rollout that has not admitted itself yet.

None of this is new

That is rather the point. Baselines before work starts, named owners, cash and non-cash kept apart. This has been standard practice in serious programme delivery for decades, and it is published, free, and written by the UK government.

AI pilots skip it almost universally. They skip it because the technology feels new enough that the old rules seem not to apply, and because taking a baseline delays the fun part by two weeks.

Those two weeks are what turns an interesting experiment into evidence. Everything else in a pilot can be recovered later. A baseline cannot.

Sources

Every figure above traces to one of these. Where a source is contested or its method has been challenged, that is said in the piece rather than left out.

  1. Teal Book, Chapter 19: Benefits management, Government Project Delivery Function · 2024

    First published 29 July 2024, updated 1 July 2026. The source of the baseline rule.

  2. HM Treasury, The Green Book: appraisal and evaluation in central government

    2020 edition with a maintenance refresh in March 2022.

  3. MIT Project NANDA, The GenAI Divide: State of AI in Business 2025 · 2025

    Method disputed. Cited here only for the buy-versus-build deployment rates.

Questions

What is a baseline in a benefits case?

The current level of performance for the thing you are trying to improve, captured before any work starts. UK government guidance requires it to be recorded on the benefits register at that point, because without it there is nothing to measure change against.

Can I count hours saved as money saved?

Not without a further decision. Hours released are a non-cash benefit. They become cash only when something changes in the budget, such as a vacancy not being backfilled or agency spend coming down. Government guidance separates cash-releasing from non-cash benefits for exactly this reason.

What if the pilot has already started and we never took a baseline?

You can sometimes reconstruct one from system timestamps, ticket logs or payroll data, provided the records predate the pilot. If you cannot, be honest that the pilot can show adoption and satisfaction but cannot show impact, and take a proper baseline before the next one.

Who should own a benefit?

A named individual with the authority or expertise to deliver it, usually the line manager whose numbers will move. Government guidance defines a benefit owner as someone who confirms the benefit's value, agrees the plan, and then takes responsibility for realising and reporting on it.


Elsewhere

The library is the other half of this: small tools built for People teams, published with the repo and the prompts. The newsletter carries one build a week.