Last week we wrote about why nearly $1.5 trillion of software spend can’t be traced to an outcome anyone can prove. The replies mostly said a version of the same thing: fine, so whose job is it to fix that?
It’s a fair question, and the honest answer is that in most organizations there are no formal roles for this. It falls between the person who decided the work mattered and the person who delivered it. In the gap between them nobody owns the sentence “here is what this was supposed to change, and here is the evidence it did.”
That gap is where the AI Product Builder comes in.
We’re not the only ones noticing
You’ve probably seen the other labels and several people are describing the same shift right now under different names.
FourWeekMBA has written about the “Builder PM”: the role that appears when AI rewrites what product management means, someone whose “spec is the prototype” and who builds specifications directly with agentic tooling rather than handing them off. Others use “AI Product Engineer.” And “AI Product Manager” is now a mature category with its own hiring guides and salary bands.
So the market has noticed that something changed. What the existing labels tend to describe, though, is a person who builds AI products, someone fluent in models, evals, agentic workflows and inference costs, pointed at an AI-powered feature. And at the same time can interface with the customer, has a deep understanding of architecture and is business savvy at the same time.
That’s a real and valuable job. It just isn’t the gap we started with anymore.
Being excellent at building AI features doesn’t, by itself, mean anyone has defined which outcome was worth pursuing or verified that it happened. You can staff a team entirely with skilled AI product managers and still produce the situation from last week’s piece: enormous spend, real output, no provable result.
The AI Product Builder is our name for the part the other labels leave out, owning the definition of the outcome and the proof it landed. Whether you end up calling that an AIPB, a Builder-PM, or just a very good product person matters less than whether someone in your organization is accountable for it.
Why the role appeared now
Roles don’t emerge because someone writes a blog post naming them. They emerge when a constraint moves.
For most of software history, building was the bottleneck. That single fact organized everything: roadmaps rationed engineering time, prioritization frameworks existed to decide what to cut, and slow building capped how much you could get wrong. You could only get so much wrong per quarter.
Then generation got cheap. Faros AI, analyzing telemetry from more than 10,000 developers across 1,255 teams, found that teams with high AI adoption complete about 21% more tasks and merge 98% more pull requests. Its larger 2026 study, covering some 22,000 developers, found epics completed per developer up 66%.
Look at what happened next, though, because this is the whole argument.
In the same Faros dataset, pull request review time rose 91%, and PR size grew 154%. LinearB’s 2026 benchmarks, drawn from 8.1 million pull requests across 4,800 organizations, found AI-generated PRs wait 4.6 times longer before a reviewer picks them up. In all fairness, LinearB also found reviewers move about twice as fast through them once they start; the delay is in getting attention, not in the reading.
And they warrant the attention. CodeRabbit’s analysis of 470 open-source pull requests found AI-generated ones carried roughly 1.7× more issues overall, with logic and correctness problems about 75% more common than in human-written PRs. (CodeRabbit sells AI code review, so read the framing with that in mind, but the direction is consistent with what the throughput data implies.)
The bottleneck simply moved, from producing work to deciding whether the work is right and taking responsibility for it.
That is a different skill from building, and organizations have not staffed for it.
The productivity illusion, and why we’re citing it carefully
There’s a finding here that ought to unsettle anyone measuring AI success by how fast their teams feel.
In early 2025, METR ran a randomized controlled trial with 16 experienced open-source developers across 246 tasks. These weren’t toy problems: developers worked in their own mature repositories, on codebases they had averaged around five years with. The developers using AI took 19% longer. Afterwards, having just lived through it, they estimated AI had made them roughly 20% faster.
A 39-point gap between what happened and what the people it happened to believed.
Now the part that matters just as much. In February 2026, METR published an update saying it is revising the experiment’s design, and that based on conversations with participants it believes developers are likely more sped up by AI tools now than that early-2025 estimate suggests. The study ran on tooling a generation behind what’s shipping today.
We’re citing it anyway, with the caveat attached, because the durable finding is the gap between measurement and perception. That gap doesn’t vamoosh when the models improve.
Set it against the most-cited study in the other direction, Peng et al. found developers completed a task 55.8% faster with GitHub Copilot, and notice that study’s boundaries too. A single synthetic exercise (implement an HTTP server), recruited developers rather than production teams, a confidence interval running from 21% to 89%. And a publication date of early 2023.
The sensible conclusion is that AI productivity is highly conditional: on task type, seniority and codebase familiarity, and that people cannot reliably perceive which case they’re in.
Which makes self-reported productivity close to worthless as evidence. If your AI programme’s business case rests on teams saying they feel faster, you have a feeling, not a finding.
Somebody has to close that gap with actual measurement. That somebody is the role we’re describing.
What an AI Product Builder actually does
Five things, in order. The order matters more than the list.
- Starts with the pain, not the prompt. The opening question isn’t “what should I ask the model?” It’s “what’s broken, who is currently paying for it to stay broken, and how much is that costing?” This is discovery, unglamorous, and the step that gets squished first when a deadline tightens.
- Agrees the measure of value before the work starts. Names the number that would show the pain is gone, and agrees it with whoever owns the budget, before anyone builds. If nobody can state what success looks like, that isn’t a scoping detail to resolve later. That is the finding, and it should stop the project.
- Assembles context deliberately. Teams blame model quality for a great deal that is really a context failure. The agent answered the question you gave it; it just wasn’t your question. Getting the right information, constraints and examples in front of an agent is engineering work, and it accounts for much of the quality difference.
- Verifies output against the outcome, not against the ticket. “The agent completed the task” and “the outcome was achieved” are different claims, and only one is worth paying for. In practice: observing what the agent actually produced, running both human and LLM review over it, and applying safety rubrics before anything ships. Given the defect-rate data, this is not optional hygiene.
- Commits to it in writing. Captures the whole thing as an outcome contract, what changes, how we’ll know, what it’s worth. This is the step that makes the previous four durable, because it removes the option of quietly redefining success afterwards.
What this means if you’re hiring, or being hired
Product hiring has been shifting toward evidence of thinking over evidence of tenure for a while now. Work samples, teardowns, published writing, recorded talks, visible decision-making frameworks, these carry weight that a résumé bullet doesn’t, because they show the reasoning rather than the destination.
We’d add one caveat we’d want applied to our own claims: a lot of the widely shared statistics about this shift trace back to vendor lead-magnets and SEO content rather than real survey work, so treat any confident percentage you see, including in our inbox, with suspicion. What’s defensible is the direction, and the direction is consistent across everyone actually doing the hiring.
For anyone interviewing right now, particularly if you’re between roles, the practical implication is direct.
Most candidates describe what they shipped and hiring managers hear it dozens of times a week. Try describing what you have change instead:
Here’s the business pain I think you have. Here’s the metric that would show it’s fixed. Here’s roughly how I’d get there, and here’s how you’d know within a certain timeframe if I was wrong.
Four sentences. They work not because they’re a clever framing device but because they are a work sample, delivered in the room, for that company’s actual business, unprompted.
And they demonstrate the one thing that is genuinely scarce. Knowing what’s worth building, and being willing to state in advance how you’d be proven wrong, is not.
If you’re on the hiring side, the same logic inverts into a useful filter. Ask a candidate what metric would have told them their last project was failing. The answers separate people quickly, and they separate them on something that predicts performance in a company that looks nothing like their last one.
The uncomfortable part
None of this is free, and it’s worth being honest about why the role is rare rather than pretending it’s obvious.
Defining outcomes in advance means committing to a number in front of people who will remember it. Verifying output honestly means sometimes reporting that expensive work didn’t achieve what it was meant to. Both are professionally uncomfortable in a way that shipping features is not, which is precisely why organizations drift toward measuring output.
The teams that get past it make the alternative more expensive, usually by writing the outcome down early enough that nobody can quietly move it.
Where this leaves us
Building has stopped being the constraint. Most organizations haven’t reorganized around that, and the engineering data shows it: far more code merged, review times up sharply, defect rates higher, and self-reported productivity that measurably doesn’t match what gets measured.
Somebody has to own the definition of the outcome and the proof it was reached. Call that person an AI Product Builder, a Builder-PM, or nothing at all, but if you can’t name who holds it in your organization, that’s the finding.
Discovery was always the bottleneck and it still is.
References
- CodeRabbit. (2025, December). *State of AI vs human code generation* [Analysis of 470 open-source pull requests]. https://www.coderabbit.ai/
- Faros AI. (2025). *The AI productivity paradox* [Telemetry from 10,000+ developers across 1,255 teams]. https://www.faros.ai/
- Faros AI. (2026). *AI engineering report 2026* [22,000 developers, 4,000+ teams]. https://www.faros.ai/
- FourWeekMBA. (2026). *The Builder-PM: The new role at the center of AI-era product organizations*. https://fourweekmba.com/ai-builder-pm-role-product-management-2026/
- LinearB. (2026). *2026 software engineering benchmarks report* [8.1 million pull requests, 4,800 organizations, 42 countries]. https://linearb.io/
- METR. (2025, July). *Measuring the impact of early-2025 AI on experienced open-source developer productivity* (arXiv:2507.09089). https://arxiv.org/abs/2507.09089
- METR. (2026, February 24). *We are changing our developer productivity experiment design*. https://metr.org/
- Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). *The impact of AI on developer productivity: Evidence from GitHub Copilot* (arXiv:2302.06590). https://arxiv.org/abs/2302.06590
—
A note on the productivity figures: the METR and GitHub Copilot studies reach opposite conclusions because they measure different things. Copilot’s 55.8% came from one synthetic task with recruited developers in 2023, with a confidence interval spanning 21% to 89%. METR’s 19% slowdown came from experienced developers working in repositories they knew deeply, using early-2025 tooling, and METR has since said the effect is likely smaller today. Both are cited deliberately. The point is not that either is definitive, but that the honest range is wide enough to make self-reported productivity unusable as evidence.
Rezoomex helps teams define the why, verify what agents actually produce, and tie payment to outcomes that can be proven.
Meeting us in person? We'll be at the Chief Product Officer Summit in San Francisco on September 24 and at Mind the Product Chicago (formerly called INDUSTRY) on October 6–7.
Want the rest of this series?
We're publishing weekly through August and September on the AI Product Builder role, agent output verification, and how payment is changing in the agentic economy.
Weekly, while the series runs. No spam, unsubscribe anytime.


