At some point the board asks the only question that matters: so what did the AI actually get us? The honest answer is usually not a line anyone can point to in the accounts. The pilot worked. The tool does what it said. People are using it. And yet no one can say, with a straight face, that the business is measurably better off. That gap — between an AI programme that is plainly working and a business that has plainly not moved — is where good programmes quietly lose their funding. Not because they failed, but because their value was never made visible on a report anyone reads.
Value delivered is not value realised. A business demands an improved business outcome, and it is indifferent to whether the lever was AI or anything else. Measurement is how you tell whether you got the outcome — and, when you didn't, where the value went instead.
This is not a soft point about being rigorous. It is the difference between keeping the programme that was working and cutting it.
You cannot prove a slope from a single point
Start with the precondition, because most of the argument follows from it. If you have not measured your current performance — or have not measured it for some time, or measured it a different way — you cannot make an objective claim about what AI changed. There is nothing to compare against. You are left asserting an improvement, and an assertion is exactly what a sceptical CFO is right to discount.
A baseline is not a formality. It is the earlier of the two points you need to draw a line. The awkward part is that the value often is real before it is visible. When a general-purpose technology arrives, the measured numbers can understate the gain for years, because the complementary work — the process redesign, the reskilling, the plumbing — is itself an investment that depresses measured output before the benefit is harvested. Economists have a name for the shape: the productivity J-curve. It is the same lag Solow pointed at in 1987 when he said you could see the computer age everywhere but in the productivity statistics. The lesson is not to be patient for its own sake. It is that if you have no baseline and no instrument, you will read the bottom of the J as failure and cut the thing on the eve of it paying off.
There is a subtler baseline error, and it is the one that flatters AI most. If you do the sensible prerequisite work first — clean the master data, fix the broken process, retire the duplicate steps — and only then switch on the AI, you must take the baseline after that work has settled. Baseline before it, and every gain from the cleanup is silently credited to the AI. That inflates the return, and it sets a false expectation for the next project that will not survive contact with reality. The enabling work is often where a significant part of the benefit comes from, which is the whole subject of another article.
Measure the numbers you already trust
The temptation, once a shiny programme is underway, is to invent shiny new metrics for it. Resist it. The fairest judge of whether AI helped is a number you already report and already believe, because you cannot quietly redefine it to flatter the result.
Financial performance measurement is well established, and the recipe is not mysterious: decompose the financial outcome into the operational drivers beneath it, and those into the process metrics beneath them. Cash resolves to days sales outstanding, which resolves to collection-cycle time. Margin resolves to cost-to-serve, which resolves to process cost per transaction. Revenue resolves to win rate and throughput, which resolve to response time. If AI is genuinely improving the operation, it should move a process metric, which should roll up to a driver, which should roll up to the financial. So run the same measurements you already run, with the same methodology, and see whether it does. Changing the ruler when you want to prove an improvement is not objective, unbiased measurement.
Keeping the existing metric has a second virtue beyond fairness: it resists gaming. Set a bespoke target for a programme and people optimise the target rather than the outcome it was meant to stand for, sharpened by Marilyn Strathern into the line everyone quotes: when a measure becomes a target, it ceases to be a good measure. A number you already report, for reasons that predate the AI, is much harder to bend.
There is one honest exception. Sometimes AI does something genuinely new that your existing metrics were never built to capture, and holding to "the same numbers" would under-credit it. But allow a new measure only if it is declared before you see the results, so it cannot become a post-hoc excuse, and only if it is relevant to business performance or influences a decision someone takes. That second test is the important one. Model accuracy, tokens processed, queries handled — these feel like measurement and inform no fund-or-cut decision. They are model telemetry, not operational business process performance. The entire reason to measure is to decide whether to keep paying for the thing, so a metric that drives no decision is noise dressed as rigour. The pull toward whatever is easy to count is not hypothetical: MIT's 2025 study of enterprise AI found budgets skewed heavily to sales and marketing, in part because those outcomes are easier to attribute than the back-office ones where the returns were actually better.
None of this is about AI
It is worth saying plainly, early, because it is what makes the rest trustworthy: none of this discipline is specific to AI. It is how you would measure the effect of any intervention — a reorganisation, a new system, a lean programme, a change of supplier. AI is simply the intervention everyone is spending on, or considering spending on, right now. There is also a competitive pressure to be seen to be adopting AI — and in that rush, the rigorous methodology is the first thing dropped, and the AI metrics quietly become the target.
Why abandoned? Because AI tempts you to skip it in a way a new forklift does not. The demonstration looks obviously impressive, so the improvement feels self-evident and beneath proof. The vendor and the board are both pushing in the same direction, so scepticism reads as obstruction. And the output is intangible and probabilistic, which makes it genuinely harder to attribute than a machine that visibly stamps more parts per hour. This is also where culture tells: organisations that empower their contrarians run a lower risk of abandoning honest measurement. The contrarian truth is that AI does not need a special kind of measurement. It needs the ordinary kind that people quietly drop the moment the technology gets exciting.
Cause is a chain, and it breaks link by link
"We deployed AI and the number went up" is a correlation, and a correlation is not what a CFO should accept. Establishing that AI caused an improvement means treating the claim as what it is: a chain of linked assertions, each of which can break on its own.
AI caused the task to improve — that link you can actually prove, with a holdout or a staggered rollout, one team or region or product line getting the tool while a comparable one does not, and the two compared. The task improvement should cause the operational driver to move — that link is a theory about the process, and it is testable. The driver should cause the financial to move — that link is close to an accounting identity, and it is testable too. Where you can isolate, isolate. Where you genuinely cannot freeze the rest of the business for two quarters, the fallback is discipline, not silence: log every concurrent change, and refuse to claim what you cannot separate from it. The reason to measure at every level rather than only at the ends is that where the chain breaks is where the constraint is. A gain that appears at the task and vanishes before the financial is not a failure to be buried. It is a finding, and it points at something.
Measure the whole stream, not its ends
This is the part most "measure your AI ROI" advice misses, and it is the part that matters most. When AI eases a bottleneck, the improvement is real at the task — but easing a constraint does not remove it, it moves it to any of the subsequent steps in the chain. Measure only at the system level and you may see nothing, or worse, a system that has got slower. Measure at the task level as well and the picture resolves: the original problem was solved, and the gain is now being eaten by a new constraint that has surfaced downstream. The gap between the two levels is the whole diagnostic.
The counterintuitive version is that making one step faster can make the whole system slower. Push more, faster output into a downstream step of fixed capacity and work-in-progress piles up in front of it, lead times stretch, and error rates climb as people rush or thrash — the plain arithmetic of Little's Law, where the work in a system is its throughput times the time each item spends in it. Speed up the analysis and delivery gets worse is not a paradox; it is what happens when you improve a step that was never the constraint, or flood the step that is.
Two ways the gain gets trapped:
- The action never happens. AI hands over faster, decision-ready output, complete with the insight to act on it — and the people downstream have no spare capacity to follow up. You automated the analysis, not the action, and value only lands when the action completes. Measure to the point of realised outcome — the invoice actually collected, not the collection recommendation generated — and watch the queue building at the human step.
- The next input is dirty. AI output feeds a process that merges it with other data of poor quality, and the bad data negates the benefit. The real constraint was never the step AI improved; it was data quality two steps over.
The instrument for seeing all of this already exists, and it is not new. It is value stream mapping — the lean practice of drawing the whole flow, value-adding steps and non-value-adding steps alike, and finding where value leaks. It pairs exactly with the theory of constraints, which tells you that the only way to lift the throughput of a system is to relax its binding constraint, and that the constraint always moves when you do. The theory tells you the constraint shifts; the map shows you where it went, because a shifted constraint reveals itself physically — as the work-in-progress and the wait time now piling up in front of the next step. One pragmatic guard so no one drowns in dashboards: map the whole stream once to find where value leaks, then permanently instrument only the few points that matter — the eased step, the emergent constraint, and the end outcome.
The honest report
Put those together and the honest way to report a trapped gain writes itself: AI delivered the task improvement it promised; the net effect at the system is currently near zero because a new constraint now binds; realising the value requires fixing that constraint. That sentence is true, it is fair to the AI, and — unlike "the AI underdelivered" — it tells you the next move.
It also protects the person most tempted to report something less honest. An AI programme manager is rewarded for showing that AI delivered value: adoption is up, the task is faster, the model performs. Every one of those can be true while the business outcome does not move — Goodhart's law, one level up, where the proxy is now "AI delivered value" and the outcome it was meant to stand for is "the business is better off." The good news is that the honest manager should want this measurement, not fear it, because it is exactly what clears them. It proves the AI did its job and names the downstream constraint eating the benefit, so when the P&L does not move they are not carrying the blame for a failure that lives two steps away in someone else's function.
Someone has to own the outcome
Which exposes the catch. All of this only pays off if someone owns the outcome from end to end. Measure AI to the task, let the constraint move downstream into another team's remit, and the constraint gets identified and then nobody fixes it — because it is, technically, not anyone's job. The measurement finds the leak; an owner of the whole value stream is what stops the leak. That is an organisational question as much as a measurement one.
The unglamorous conclusion
The business, at the system level, demands an improved business outcome. That objective does not care whether the lever was AI, a reorganisation, or a better supplier — and it is not satisfied by "the intervention delivered value" if the value never reaches the outcome. Understand how the operational tasks depend on one another, measure the whole chain methodically and with the ruler you already trust, establish cause link by link, find the constraint that is suppressing the benefit, and fix it. The measurement approach is what lets you see all of that at once.
AI is the occasion for this discipline, not the subject of it. The reason to insist on it now is not that AI is uniquely suspect, but that it is uniquely good at looking like progress. And a business that mistakes value delivered for value realised will keep funding the programmes that impress it and cut the ones that were quietly, invisibly, about to pay.
Sources
- Erik Brynjolfsson, Daniel Rock and Chad Syverson, "The Productivity J-Curve: How Intangibles Complement General Purpose Technologies", American Economic Journal: Macroeconomics 13, no. 1 (2021): 333–372 (NBER Working Paper 25148, 2018) — intangible complementary investment causes measured productivity to be understated in a general-purpose technology's early years and overstated later. https://www.nber.org/papers/w25148
- Robert M. Solow, "We'd Better Watch Out" (review of Manufacturing Matters by Stephen S. Cohen and John Zysman), New York Times Book Review, 12 July 1987 — "You can see the computer age everywhere but in the productivity statistics."
- Charles Goodhart (1975); the widely quoted formulation is Marilyn Strathern's (1997): "When a measure becomes a target, it ceases to be a good measure." Strathern, "'Improving ratings': audit in the British University system", European Review 5, no. 3 (1997): 305–321.
- John D. C. Little, "A Proof for the Queuing Formula: L = λW", Operations Research 9, no. 3 (1961): 383–387 — the average work-in-progress in a stable system equals throughput multiplied by the average time each item spends in it (WIP = throughput × lead time).
- Mike Rother and John Shook, Learning to See: Value-Stream Mapping to Add Value and Eliminate Muda (Lean Enterprise Institute, 1998) — the lean practice of mapping an entire flow, value-adding and non-value-adding steps alike, to locate where value is lost.
- Eliyahu M. Goldratt and Jeff Cox, The Goal: A Process of Ongoing Improvement (North River Press, 1984) — the theory of constraints: system throughput is set by the binding constraint, so only elevating that constraint improves the system, and the constraint moves to the next step once it is relaxed.
- MIT NANDA (Project NANDA), The GenAI Divide: State of AI in Business 2025 (July 2025) — enterprise GenAI budgets skewed heavily to sales and marketing even though back-office automation showed better returns, in part because front-office outcomes are easier to attribute and measure. https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
Ari Das-Purkayastha advises organisations on delivering and managing technology-led transformation.