It’s budget season. Somewhere in your organization, a slide is being built that puts a return next to last year’s AI spend.
Before it reaches the CFO, it’s worth asking a plain question about that number: where did it come from? Not which model produced it, but who measured what, against what baseline, agreed when?
For most organizations the honest answer is that nobody measured anything. The number is a reconstruction, assembled after the fact from the initiatives that happened to go well. That is not a failure of diligence. It is the predictable result of a decision made much earlier, when the work started without a success metric anyone wrote down.
We went through what the research actually says about enterprise AI ROI. The picture is more useful than the one in circulation, and considerably less flattering.
The AI ROI statistic everyone quotes is the weakest one
You have seen it. Ninety-five percent of enterprise AI pilots deliver no measurable return. It has been on conference slides since August 2025, and it will be on several at the summits this autumn.
Start with what the source actually says. The claim comes from The GenAI Divide: State of AI in Business 2025, published in July 2025 by Project NANDA, a research initiative inside the MIT Media Lab. Its wording is that 95% of organizations are getting zero return, not that 95% of pilots produce nothing measurable, which is the form now in circulation. Those are different claims about different units.
The report is not peer reviewed, but it does disclose its evidence base, and more than it usually gets credit for: structured interviews with representatives of 52 organizations, survey responses from 153 senior leaders collected across four industry conferences, and a review of more than 300 publicly disclosed AI initiatives. The weakness isn’t secrecy. It’s that 153 people standing at industry conferences are not a random draw from the enterprise population, and the report generalizes as though they were.
The harder problem is the headline number itself. Wharton’s Kevin Werbach, in a public post picked up by Futuriom, wrote that he had read the document several times and still could not work out where the 95% comes from. He noted that a five-percent figure does appear in section 3.2, describing custom enterprise AI tools that were successfully implemented, a far narrower claim than the one the slides make, and not obviously the source of the headline either.
So the most-repeated number in enterprise AI is an unreviewed figure with no visible derivation, restated in a stronger form than its own source used, circulating between decks that cite each other.
What the larger surveys actually show
There is better evidence, and it is not good news.
McKinsey’s State of AI survey, published August 2026, ran across 1,719 participants in 97 countries between 4 May and 8 June 2026, with 36% of respondents at organizations above $1 billion in revenue. Two findings matter for a budget conversation.
37% of respondents attribute at least some EBIT impact to AI. That figure is essentially unchanged from a year earlier.
6% qualify as high performers, meaning organizations attributing 5% or more of EBIT to AI use and describing that impact as significant.
Read those two numbers together with a third: 40% of large organizations ($1 billion or more in revenue) now report scaling AI agents, up from 27%. Deployment is accelerating. Reported earnings impact is flat.
That combination rules out the most comfortable explanation. If the problem were that adoption hadn’t reached scale yet, scaling would be moving the EBIT line, but it isn’t.
The gap that should worry a CPO
Here is the finding that is truly noteworthy.
In the same survey, 80% of respondents say AI has improved their individual productivity. 37% can point to any enterprise earnings impact at all. Six percent can point to a material one.
Eighty percent of people experiencing a personal improvement, and a flat line where the money is, is not a contradiction that resolves itself with more deployment. It is what happens when time is saved in places where nobody had committed to converting saved time into anything. The engineer finishes the ticket ninety minutes early. The analyst drafts the memo in one pass instead of three. Both are real. Neither shows up in EBIT unless someone decided, in advance, what that recovered capacity was for and how they would know it had been redeployed.
Note also what “at least some EBIT impact” is doing in that 37%. The Register, covering the survey, remarked that the phrase was doing a lot of heavy lifting, and that is the right instinct. It is a self-report, on a survey, with no instrument behind it. An executive who cannot measure will not answer “none.” They will answer “some.” The most likely reading of the 37% is that it represents belief, and the 6% represents measurement.
“Unclear business value” is not a technology problem
Gartner’s prediction for agentic projects points at the same place from a different angle: over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value or inadequate risk controls.
Two of those three reasons are measurement failures wearing different clothes.
“Unclear business value” does not mean the value was absent. It means nobody could produce it on demand when the project came up for review. A project with real but unevidenced value and a project with no value look identical in a budget meeting, and they die the same death.
“Escalating costs” is the same story from the other side. Cost escalates without alarm when there is no denominator. A team that agreed up front that the outcome was worth, say, $400,000 in reduced manual reconciliation notices at $300,000 of token spend that something has gone wrong. A team that agreed nothing notices at cancellation.
Gartner’s analysis adds a detail worth carrying into vendor conversations this quarter: of the thousands of vendors positioning themselves as agentic, the firm considers roughly 130 to be genuinely so. Senior Director Analyst Anushree Verma’s assessment was that most agentic AI projects at present are “early stage experiments or proof of concepts that are mostly driven by hype and are often misapplied.”
What separates the 6%
McKinsey’s survey does say what distinguishes its high performers, and the answers are unglamorous.
They redesign workflows because of AI rather than layering AI onto existing ones: nearly three-quarters of high performers report fundamentally redesigning workflows because of their AI use, up from 55% a year earlier, against roughly one-quarter of everyone else. They pursue growth and innovation alongside efficiency, rather than efficiency alone. Their senior leaders demonstrate commitment at twice the rate. And they are twice as likely as other respondents to have defined processes for quantifying the impact of their AI initiatives.
That last one is not a virtue that happens to correlate with the others. It is the mechanism. An organization with a defined measurement process is an organization that had to name the outcome before starting, because you cannot design a measurement for a goal you haven’t stated. The workflow redesign follows from the same discipline: you only rebuild a process when you know which number the rebuild is supposed to move.
The 6% are not better at AI. They are better at deciding what they wanted before they bought anything.
The honest caveat
Everything above is survey data. That is a real limitation and it should be stated plainly rather than buried.
Nobody has run a controlled study of enterprise-level AI ROI. Randomised experiments on AI and productivity do exist and some are good, on individual task performance, on consultant output, on developer speed, but none of them measure what AI did to a quarter’s EBIT. At the enterprise level there is no randomization, no control group, no independent verification of the figures.
McKinsey’s numbers are what executives report about their own organizations, and the same self-report problem that inflates the 37% could be distorting the 6% in either direction. The Gartner figure is a forecast, not an observation. The firm-level consulting claims circulating about agentic delivery, including McKinsey’s own May 2026 finding of threefold to fivefold productivity improvements with a 60% reduction in team size, come from observation of client engagements rather than from studies with methodology sections.
So the defensible position is narrower than the headlines: reported enterprise-level returns from AI are low and flat, the organizations reporting real returns are distinguished mainly by having defined how they would measure them, and no one has established this causally. Anyone quoting you a precise industry ROI percentage is quoting a survey average or an invention.
The absence of a clean number is itself the finding. If enterprise AI returns were straightforwardly measurable, somebody would have measured them by now.
What to do before the budget goes in
For every AI initiative in next year’s ask, fill in four lines.
- The business pain, in the business’s words. Not “adopt AI in support.” The pain: resolution time on tier-two tickets averages 31 hours and enterprise renewals cite it. If you cannot write the pain without naming a technology, you have a solution looking for a problem.
- The metric that proves it’s fixed. One number, already known to the business, that someone outside the project team already watches. New metrics invented for the initiative are a warning sign. They tend to be the ones that can be moved.
- Today’s baseline. This is the line that gets skipped, and it is the only one that expires. A baseline can only be taken before the work starts. Every initiative already underway without one has permanently lost the ability to prove its own result, and the reconstruction will be an estimate no matter how good the analyst is.
- Who verifies it, and when. A named person outside the delivery team, and a date. Verification that the delivery team performs on itself is not verification. It is a status update.
An initiative with all four is a budget line you can defend under questioning. An initiative missing line three is an anecdote with a cost attached, which is exactly what the 40% cancellation forecast is describing.
It is worth running this over your existing portfolio before you build next year’s. Our experience is that the exercise is uncomfortable and fast: most teams can tell you the pain and the metric within a minute, and go quiet at the baseline.
The short version
The 95% statistic comes from an unreviewed report with no shown derivation, and the popular version is stronger than what the source actually says. Stop citing it.
Better evidence is worse news. 37% of organizations report any EBIT impact from AI, flat year on year, while agent scaling has climbed from 27% to 40% of large enterprises.
80% report personal productivity gains against 37% reporting any earnings impact. Saved time does not convert itself.
Gartner expects over 40% of agentic projects to be canceled by end of 2027, citing escalating costs and unclear business value, both measurement failures before they are technology failures.
The 6% seeing real returns are distinguished by having defined processes to measure impact. That is a decision made before the work, not a capability bought during it.
Before budget season: pain, metric, baseline, verifier. The baseline is the one you cannot go back for.
You cannot compute a return on an outcome you never defined. Most AI ROI is lost before a single agent runs.
We’ll be at the Chief Product Officer Summit in San Francisco on Thursday 24 September, showing exactly this: an outcome defined in business terms, captured as a contract, and an agent’s output verified against it, live, on real work. If you’re building next year’s AI budget, it’s a useful twenty minutes.
Book a time at the CPO Summit →
Next in this series: the outcome contract itself: what goes in one, and how the verified result becomes the thing you pay against.
References
- Gartner. (2025, June 25). Gartner predicts over 40% of agentic AI projects will be canceled by end of 2027 [Press release]. gartner.com
- McKinsey & Company. (2026, August 25). The state of AI in 2026: On the road to ROI. QuantumBlack. n = 1,719 across 97 nations, fielded 4 May–8 June 2026. Dated PDF, because the /the-state-of-ai landing page rolls over each year. mckinsey.com (PDF)
- McKinsey & Company. (2026, May 28). Rewiring software delivery for the agentic era. mckinsey.com
- Challapally, A., Pease, C., Raskar, R., & Chari, P. (2025, July). The GenAI divide: State of AI in business 2025. Project NANDA, MIT Media Lab. Not peer reviewed. Stated basis: structured interviews with 52 organizations, 153 survey responses collected at four industry conferences, and review of 300+ publicly disclosed AI initiatives. media.mit.edu
- Futuriom. (2025, August 26). Why we don’t believe MIT NANDA’s weird AI study. Quotes Kevin Werbach (Professor and Chair of Legal Studies & Business Ethics, Wharton) from his own public post. futuriom.com
- Vigliarolo, B. (2026, August 25). McKinsey says enterprise AI is finally “on the road to ROI.” The Register. theregister.com
A note on sources. None of the evidence here is experimental. The McKinsey figures are the strongest available because the sample is large, international and consistently fielded year over year, but they remain executive self-reports about their own organizations, which is exactly the weakness the article argues about. The Gartner cancellation figure is a forecast and should be read as an analyst’s judgement, not a measurement. The NANDA report is the weakest source cited and appears here only as the subject of a correction. Consulting claims about productivity multiples come from client observation and carry no methodology section; we have flagged them as such rather than dropping them, because they are widely quoted and readers will meet them elsewhere. Where these disagree, trust them in that order. During fact-checking we corrected four of our own claims: we had mis-stated the NANDA report’s methodology, called it an MIT spinoff when it is a project inside the Media Lab, quoted a workflow-redesign figure to a precision the source does not use, and claimed no controlled studies of AI productivity exist when what is missing is controlled studies at the enterprise-EBIT level. The irony of getting sourcing wrong in a piece about sourcing is not lost on us, and it is the reason this note exists.


