This year, the world will spend nearly $1.5 trillion on software. Not on IT overall, that figure is $6.37 trillion, just on software. It is a 15.5% jump over 2025, and it is one of the fastest-growing lines in any enterprise budget. It is also why enterprise AI ROI has become the hardest question a product leader can be asked to answer.
Here is the uncomfortable question that goes with it: how much of that spend can be traced to a business outcome someone can prove?
Not “was it delivered.” Not “did it ship on time.” Not “are people logging in.” Proved, as in, there was a business pain, we agreed on the metric that would show it was fixed, and that metric moved.
For most organizations, the honest answer is that nobody knows. And the evidence that has accumulated over the past eighteen months suggests the answer, where it can be measured at all, is worse than anyone expected.
What the enterprise AI ROI data actually shows
The first is the one that made headlines. MIT Media Lab’s Project NANDA, in a widely circulated 2025 report titled The GenAI Divide, reviewed more than 300 publicly disclosed AI initiatives alongside structured interviews at 52 organizations and survey responses from 153 senior leaders. Its headline finding: roughly 95% of enterprise GenAI pilots produced no measurable return on P&L. Only about 5% delivered rapid, material impact. The report put enterprise investment behind those pilots at $30-40 billion.
The study is worth citing carefully, it is labelled preliminary findings and has not been peer-reviewed, and its 95% figure has been contested. But the direction it points is corroborated by everything that follows.
The detail that matters most is buried in the diagnosis. MIT did not conclude that the models were inadequate. It attributed the failure to what it called a learning gap, an organizational inability to fit AI into real workflows, structures, and decisions. The technology worked. The wrapping around it did not.
The second comes from Gartner, which predicts that more than 40% of agentic AI projects will be canceled by the end of 2027. The reasons it gives are worth reading slowly: escalating costs, unclear business value, or inadequate risk controls.
“Most agentic AI propositions lack significant value or return on investment (ROI), as current models don’t have the maturity and agency to autonomously achieve complex business goals or follow nuanced instructions over time.”
– Anushree Verma, Senior Director Analyst, Gartner
Gartner also estimates that of the thousands of vendors now marketing “agentic AI,” only about 130 are real, the rest are engaged in what it bluntly calls agent washing: rebranding assistants, chatbots, and RPA that existed before.
The third is quieter but cuts the deepest. In a Gartner survey of 204 finance leaders, 45% said their AI investments lean toward productivity, while only 20% lean toward decision quality. The returns ran the other way: only 17% reported significant or transformational value from productivity-focused investments, against 31% for decision-quality initiatives. Organizations are concentrating spend in the category that pays back least, and, notably, in the category that is hardest to attribute to any single business result.
Put those together and a pattern emerges. Spend is accelerating. Confidence is not. The enterprise AI ROI problem is real and measurable, and the gap between the two is not a technology gap.
The comfortable explanations, and why they’re wrong
When enterprise AI ROI disappoints, four explanations usually get offered. Each is plausible but none survives what the data says.
“The models aren’t good enough yet.” Model capability has improved dramatically over the past two years, and benchmark performance has climbed steadily. Yet Gartner’s cancellation forecast attributes failures to costs, unclear value, and weak risk controls, not to model capability. When the reasons projects die are commercial and organizational, better models do not address them.
“We need better tooling.” Tooling has never been more abundant or cheaper. Access to AI was never the bottleneck.
“It’s a change management problem.” Closer, this is essentially MIT’s “learning gap”, but it describes the symptom rather than the cause. Adoption stalls for a reason. People do not resist tools that visibly solve a problem they have.
“We need more data.” Often true, and almost never the binding constraint. Plenty of well-instrumented organizations produce beautifully measured features that nobody wanted.
There is a simpler explanation, and it predates AI by decades.
The actual diagnosis: everyone defines the what, almost nobody defines the why
Enterprise teams are genuinely excellent at specifying the what. The spec. The PRD. The acceptance criteria. The sprint scope. This is a mature, well-tooled discipline with thirty years of practice behind it.
The why is a different question: what business pain is acute enough that someone would pay to solve it, and what metric would prove it was solved?
That question is asked far less often than anyone would like to admit. It is frequently assumed, gestured at in a strategy deck, or delegated upward to somebody presumed to have settled it already. And when the why is undefined, everything downstream still works, which is precisely what makes it so dangerous. Requirements get written. Sprints run. Software ships. Dashboards go green.
The result is a system that executes flawlessly against an unexamined premise.
Pendo’s 2019 Feature Adoption Report put a number on the cost of this years before agents entered the picture: roughly 80% of features in the average software product are rarely or never used, with about 12% of features generating 80% of daily usage volume. Those unused features were not built by incompetent teams. They were built by competent teams executing a well-specified what that was never tied to a verified why.
Marty Cagan has spent years making the same argument to product organizations, distinguishing teams “measured by outcome” from feature factories measured by output. His summary of why the shift is so rare is honest about the difficulty:
“There’s no question that delivering outcomes is much more difficult than delivering output.”
– Marty Cagan, Silicon Valley Product Group
He is right, and it explains the behaviour. Output is legible, schedulable, and easy to report. Outcomes are contested, slow, and occasionally embarrassing. Given a choice between a metric that makes progress visible and one that makes truth visible, most organizations pick visibility of progress.
Discovery, the work of finding the pain worth solving and defining the proof, is the hard part. It is also the part that gets compressed first when a deadline tightens.
What agents changed, and what they didn’t
Here is where the old problem becomes an expensive new one.
For most of software history, a poorly defined why was throttled by the cost of building. Discovery was sloppy, but delivery was slow and expensive enough to act as a brake. You could only build so much of the wrong thing per quarter.
Agents removed the brake.
An agent will pursue a target with extraordinary speed and no instinct for whether the target was worth pursuing. It does not pause to ask whether the ticket reflects a real business pain. It executes the what, precisely, tirelessly, and at a cost that scales with effort rather than results. Undefined objectives now consume tokens continuously rather than sprint capacity intermittently.
This is why Gartner’s cancellation forecast lists escalating costs and unclear business value side by side. They are not two independent risks. They are the same enterprise AI ROI failure observed from the finance seat and the strategy seat. When the objective is vague, an agent’s speed converts directly into spend, and its output converts into work that looks finished without anyone able to confirm it was right.
Faster building industrialized the outcome gap instead of fixing it.
The part nobody wants to say out loud: the billing model rewards this
There is a structural reason the gap persists, and it is not incompetence.
Enterprise software is still largely bought the way it was twenty years ago, by hours, seats, sprints, and story points. Every one of those units measures effort. None measures result. When the commercial model pays for activity, activity is what gets optimized, and no amount of outcome rhetoric will overcome a contract that pays regardless.
The market has started to notice. In Futurum Research’s 1H 2026 survey of enterprise software buyers, fewer than one in five still prefer classic per-user pricing. Some 43% now prefer consumption-based models and 27% prefer outcome-based structures.
Vendors are responding where the outcome is unambiguous enough to price. Zendesk charges for successful AI-driven issue resolutions rather than AI seats. Intercom prices its Fin agent on resolutions. Decagon structures pricing around resolved customer interactions.
The logic is hard to argue with, because per-seat pricing has become structurally perverse in an agentic world: the better the agent performs, the fewer seats the customer needs, so the vendor is effectively paid to under-deliver.
But outcome-based pricing only works if you can do one thing first. You have to be able to verify the outcome. Which brings the problem back to where it started.
Fixing enterprise AI ROI is a sequence
None of this is solved by buying a better model, and none of it is solved by a pricing page. It requires doing three things in order, and the order is the whole point.
- Define the why before the what. Start with the business pain, not the feature request. Name the metric that will prove it is fixed, and agree on it with the person whose budget is at stake, before any work begins. If nobody can state what would count as success, that is not a scoping problem to resolve later. It is the finding.
- Verify the output against that definition. “The agent completed the task” and “the outcome was achieved” are different claims, and only one of them is worth paying for. Verification means assembling the right context up front, observing what the agent actually produced, running human and LLM review over it, and applying safety rubrics before anything ships. Unverified agent output is not productivity. It is faster technical debt.
- Tie the commercial terms to the verified result. Capture the commitment as an outcome contract: the defined outcome, the metric, the verification method, and payment attached to the verified result, not to hours logged and not to tokens spent. This is the step that makes the first two durable, because it removes the incentive to declare victory early.
Most organizations attempt step three without steps one and two, which is why outcome-based pricing conversations so often stall. You cannot pay for a verified outcome if you never defined the outcome or built the means to verify it.
References
Cagan, M., & Castro, F. (2025, March 17). *Outcomes are hard*. Silicon Valley Product Group. https://www.svpg.com/outcomes-are-hard/
Challapally, A., Pease, C., Raskar, R., & Chari, P. (2025). *The GenAI divide: State of AI in business 2025* [Preliminary findings report]. MIT NANDA. https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
Futurum Group. (2026, May 12). *Are outcome-based and hybrid AI pricing models rewriting the vendor playbook?* [Press release]. https://futurumgroup.com/press-release/are-outcome-based-and-hybrid-ai-pricing-models-rewriting-the-vendor-playbook/
Gartner. (2025, June 25). *Gartner predicts over 40% of agentic AI projects will be canceled by end of 2027* [Press release]. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
Gartner. (2026, July 20). *Gartner survey shows 45% of CFOs say their AI investments lean toward productivity, while 20% say these investments lean toward decision quality* [Press release]. https://www.gartner.com/en/newsroom/press-releases/2026-07-20-gartner-survey-shows-45-percent-of-cfos-say-their-ai-investments-lean-towwards-productivity-while-20-percent-say-these-investments-lean-towards-decision-quality
Gartner. (2026, July 27). *Gartner forecasts worldwide IT spending to grow 14.2% in 2026, totaling $6.37 trillion* [Press release]. https://www.gartner.com/en/newsroom/press-releases/2026-07-27-gartner-forecasts-worldwide-it-spending-to-grow-14-point-2-percent-in-2026-totaling-6-point-37-trillion
Thomas, S. (2019, February 5). *The 2019 feature adoption report*. Pendo. https://www.pendo.io/resources/the-2019-feature-adoption-report/
—
Note on the MIT NANDA report: The GenAI Divide is labelled preliminary findings and has not been peer-reviewed. Its 95% figure has been publicly contested, and the report carries a disclaimer that the views expressed do not reflect the positions of any affiliated employers. It is cited here as directional evidence, corroborated by the Gartner findings above.
Rezoomex helps teams define the why, verify what agents actually produce, and tie payment to outcomes that can be proven.
Meeting us in person? We'll be at the Chief Product Officer Summit in San Francisco on September 24 and at Mind the Product Chicago (formerly called INDUSTRY) on October 6–7.
Curious whether the AI Product Builder role fits how you already work? That's the next piece in this series.
Want the rest of this series?
We're publishing weekly through August and September on the AI Product Builder role, agent output verification, and how payment is changing in the agentic economy.
Weekly, while the series runs. No spam, unsubscribe anytime.


