The software industry spent twenty years billing for effort. Hours, then days, then story points and sprints. Different units, same idea: measure what went in, because measuring what came out was too hard.

AI was supposed to break that. Building got cheap, and any remaining connection between hours worked and value delivered snapped. This was the moment outcome-based pricing should have arrived: if a machine does the work, nobody has to pretend that hours are a proxy for value any more.

Look at what the tooling market actually did with that opportunity.

We kept billing for effort. We just made the unit smaller.

GitHub Copilot moved every plan onto AI Credits on 1 June 2026, priced at one cent each. Agentic and chat features draw down a credit balance; code completions stay outside it.

Cursor meters Teams usage in dollars against per-million-token rates, with a $40 per user monthly seat underneath it and a separate token rate applied to third-party models.

Devin restructured in April 2026 into plans running from $20 to $200, plus a Teams tier at $80 a month with a per-seat charge on top. Each plan carries a usage quota that refreshes on a schedule; exceed it and you buy more, consumed at API rates.

It’s worth being fair about the direction of travel in AI agent pricing here, because the easy version of this story is wrong. These changes were not uniformly a squeeze. GitHub’s restructure layered variable “flex” allowance on top of base credits, so a Pro seat carries $10 of base plus $5 of flex, and the top tier doubles $100 into $200 of included usage. Several vendors gave buyers more included consumption, not less. Basic completions remain unmetered almost everywhere.

So this isn’t vendors being greedy. It’s something more interesting: an industry with every commercial incentive to move to outcome-based pricing, choosing not to, and building increasingly sophisticated ways to meter effort instead.

Different mechanics, one idea. You are billed for what the agent consumes, not for what it achieves. Tokens, credits, quota, every unit measures the work the machine did, and none of them says whether the work was any good.

What that means when the agent is wrong

Here’s the consequence that doesn’t appear on any pricing page.

An agent that solves your problem on the first attempt consumes one unit of budget. An agent that misreads the context, produces something plausible and wrong, gets corrected, and tries again consumes five or six. Same outcome. Several times the bill.

Under consumption pricing, the buyer absorbs the entire cost of the agent being wrong.

That’s the same structural perversity as per-seat pricing, just faster and harder to see. Per-seat quietly rewarded a vendor for building an agent that didn’t fully replace anyone. Consumption pricing does nothing to discourage one that needs a few more attempts.

And there’s a compounding problem: most teams cannot distinguish the two cases. If nobody verifies whether an attempt achieved the outcome, then six attempts and one attempt look identical on the invoice, just a bigger number. You are paying per try, with no record of which try worked.

Buyers have noticed, though not precisely enough

This isn’t only a developer-tools story. Across enterprise software, appetite for paying by headcount has collapsed.

In Futurum Research’s 1H 2026 survey of enterprise software decision-makers, fewer than one in five buyers still prefer classic per-user pricing. Some 43% now prefer consumption-based models and 27% prefer outcome-based structures. Futurum’s blunter finding: vendors restricted to seat-only pricing “risk immediate disqualification.”

Look at the split, though. Consumption-based pricing leads at 43%, and consumption is still effort, just metered more finely. Genuine preference for outcome-based pricing sits at 27%, and it is far harder to actually buy than to prefer.

The market has largely moved from paying for the wrong thing slowly to paying for the wrong thing quickly.

One vendor is trying to fix this, and how they’re doing it is instructive

It would be easy to write this as an industry that hasn’t noticed. That would be wrong, and the exception is worth taking seriously.

In June 2026 Cognition, which makes Devin, published what it calls an AI Productivity Guarantee. An estimator agent reviews each completed session, judges whether it produced useful output, and estimates how many human engineering hours it displaced, validated against engineers’ own time estimates. Those hours are converted at a standard rate and compared against what the customer actually consumed. If the value falls short, Cognition issues credits, up to $10 million.

Their reasoning for building it could have been lifted from this series. Scott Wu, writing for Cognition, put it as: dashboards show activity metrics like tokens consumed and lines of code generated, but none of them answer how much value the business is actually getting out.

That is a real attempt to price on results, from a company with something to lose by doing it. Credit where it’s due.

It also illustrates, precisely, why this is hard. Three things about it are worth noticing:

The vendor’s agent does the measuring. An estimator built by Cognition assesses whether Cognition’s product delivered value. That may well be accurate, but it’s the supplier’s account of the supplier’s work, which is the thing verification is supposed to remove.

It measures displaced hours, not business outcome. “How many engineering hours would this have taken a human?” is a much better proxy than tokens consumed. It’s still a proxy for effort. Nobody’s revenue moved, no customer’s problem got solved, an expensive feature nobody wanted would score well if an agent built it efficiently.

It’s an enterprise arrangement. Most buyers are on self-serve plans where the meter simply runs.

None of that makes it bad. It makes it the best available answer to a genuinely hard problem, and it shows the shape of what’s missing, which is an independent way to establish that the outcome happened.

Why outcome-based pricing is hard in software specifically

Outcome-based pricing works where an outcome is countable, frequent and objective. Customer support is the obvious case: vendors there publish per-resolution rates, because a resolution is a discrete event you can count.

Even there it’s harder than it looks. Intercom’s Fin started at a flat rate per resolution and has since had to split the definition into several distinct billable events, a resolution, a handoff into a defined procedure, a disqualification, a qualification, each priced differently. That’s the easy category: high volume, short cycle, unambiguous ending. It still took four definitions to describe what “it worked” means.

Software delivery has none of those properties. The outcome is slow, contested, and attributable to several causes at once. Did retention improve because of the feature, the pricing change, or the season? Six months might pass before anyone can say, and three other things will have shipped by then.

So the honest position is this: outcome-based pricing for software delivery isn’t blocked by unwillingness. It’s blocked by unprovability.

Nobody signs a contract that pays on results when the results can’t be established. Lawyers won’t approve a term whose central definition is absent. So the deal reverts to what can be counted, seats, hours, tokens, quota, not because anyone prefers it, but because it’s the only thing both sides can agree happened.

What has to exist first

Outcome-based pricing isn’t a pricing innovation you can adopt on its own. It’s the last step in a sequence, and it fails without the earlier ones:

  1. Define the outcome in business terms, the change, not the deliverable. “Cut enterprise time-to-first-response by a third,” not “ship the routing service.”
  2. Agree the metric before the work starts, with whoever owns the budget, including the counter-metric: what must not get worse. Throughput that rises while defects rise is not a win.
  3. Specify the verification method, who checks, against what source of truth, how often, and what happens when the two sides disagree.
  4. Attach payment to the verified result.

Most organizations attempt step four having skipped step three. That’s the entire failure mode. You cannot pay for a verified outcome if nobody specified what verification means or who performs it.

Why this lands on the product person’s desk

Notice who owns each of those four steps.

Defining the outcome in business terms is product work. Agreeing the metric and the counter-metric with the budget holder is product work. Specifying how the result gets verified is product work. Only step four, attaching payment, belongs to procurement, and it’s the only one procurement can’t do without the other three.

Which is why outcome-based pricing keeps stalling in contract review. Legal is handed a document with a payment trigger and no definition behind it, and does the responsible thing: sends it back. Everyone concludes the model isn’t ready. The model was fine. The definition was missing, and the person who could have written it wasn’t in the room.

This is the practical case for what we’ve been calling the AI product builder, someone who can state an intended business change, attach a measure to it, specify how that measure gets checked, and hold that structure through delivery. Under effort-based billing, that skill was valuable. Under any pricing model tied to results, it’s load-bearing. Nothing else in the sequence works without it.

If you’re a product leader, this is the shift worth preparing for. If you’re a product person between roles, it’s a sharper thing to say in an interview than a list of features you shipped: describe an outcome you defined, the counter-metric you set, and how you proved the change was real.

What to negotiate instead of the unit price

Whether you’re buying tooling, an agency or a delivery partner, the rate per unit is the least interesting number on the page. Five terms decide whether you have outcome-based pricing or a consumption contract with aspirational language on top.

Who absorbs rework. If the agent or the team gets it wrong, whose budget pays for the second attempt? Under consumption pricing the answer is always yours. Make that explicit rather than discovering it on an invoice.

The source of truth. Whose system decides, theirs, yours, or a shared record? If measurement runs entirely inside the supplier’s platform, you are accepting their account of their own work as the basis for payment.

The definition, written out with examples. Not a phrase but a paragraph, naming what counts and what doesn’t. Does a reverted change count as delivered? A correct implementation of the wrong requirement?

The dispute process. What happens when you disagree, who reviews it, on what timescale. Without one, the definition is decorative.

The counter-metric. One number that must not degrade, so the measure can’t be gamed at the expense of the thing you actually wanted.

The uncomfortable summary

The agentic economy was supposed to end billing for effort. Mostly it produced a finer-grained version of it, and moved the cost of failure onto the buyer.

The vendors trying to escape that are running into the same wall everyone else is: effort is the only thing anyone can currently count. Outcome-based pricing isn’t waiting on a pricing consultant. It’s waiting on infrastructure for proving that something worked.

Every piece in this series has arrived here from a different direction. The spending problem, the missing role, the verification gap, and now the pricing shift are one problem, organizations are far better at producing work than at proving what it changed.

Payment is simply where that gets expensive enough to force the issue.

Don’t negotiate the price per unit until you’ve agreed who gets to say the outcome happened.

Rezoomex helps teams define outcomes in business terms, verify what agents actually produce, and tie payment to results that can be proven, in that order.

Next in this series: what we’re hearing from product leaders heading into budget season, and where the outcome conversation is landing with CPOs.

References

  1. Futurum Group. (2026, May 12). Are outcome-based and hybrid AI pricing models rewriting the vendor playbook? [Press release]. https://futurumgroup.com/press-release/are-outcome-based-and-hybrid-ai-pricing-models-rewriting-the-vendor-playbook/
  2. Wu, S. (2026, June 4). The AI productivity guarantee. Cognition. https://cognition.com/blog/ai-guarantee
  3. Cognition. (2026, April 14). New self-serve plans for Devin. https://cognition.com/blog/new-self-serve-plans-for-devin
  4. GitHub. (2026). Copilot billing: models and pricing. https://docs.github.com/en/copilot/reference/copilot-billing/models-and-pricing
  5. Cursor. (2026). Pricing and usage. https://cursor.com/docs

A note on sources. Pricing in this category changed substantially between April and June 2026, and a great deal of published comparison content is now wrong, including several widely repeated figures we removed during fact-checking after tracing them to a single unattributed blog post from 2025. Everything above comes from vendors’ own current billing documentation or announcements, or from Futurum’s independent survey. Rates move quickly; verify directly before relying on any of them in a negotiation.

Rezoomex helps AI Product Builders define the why, verify what agents actually produce, and tie payment to outcomes that can be proven.

Meeting us in person? We'll be at the Chief Product Officer Summit in San Francisco on September 24 and at Mind the Product Chicago (formerly called INDUSTRY) on October 6–7.

Want the rest of this series?

We're publishing weekly through August and September on the AI Product Builder role, agent output verification, and how payment is changing in the agentic economy.

Weekly, while the series runs. No spam, unsubscribe anytime.