
Why 95% of AI Pilots Fail to Show ROI (And How to Be in the 5%)
MIT's 2025 study of 300 enterprise AI deployments found 95% never move the P&L. Here's what the other 5% actually did differently, and what it means before you sign your next AI vendor contract.
In July 2025, MIT's Project NANDA published a study that should worry anyone about to sign an AI vendor contract. After reviewing 300 public AI deployments, surveying 153 business leaders, and interviewing 52 executives, the researchers found that 95% of generative AI pilots inside companies were producing no measurable effect on profit or loss, despite an estimated $30-40 billion in enterprise spending already committed. This wasn't a study about whether AI works. It was a study about why most companies fail to make it work, and the answer has almost nothing to do with which model they picked.
What MIT actually measured
Researchers at MIT's NANDA initiative didn't just poll opinions. Across 300 public deployments and 52 executive interviews, they separated "adoption," a team tried a tool, from "transformation," the tool changed a real P&L line. The gap between those two numbers is what the report, titled "The GenAI Divide: State of AI in Business 2025," calls out directly: nearly every mid-size and large company has piloted something by now, but as Forbes reported on the study, only 5% of those pilots ever reach the point where a finance team can point to a number on a statement and say "that came from the AI."
The real problem isn't the model, it's the learning gap
The instinct when a pilot underperforms is to blame the model, wait for the next release, and try again. MIT's researchers found that's the wrong diagnosis. Their core finding was that most of the GenAI systems that failed to move the needle "do not retain feedback, adapt to context, or improve over time." In practice, that means a chatbot gets bolted onto a support queue or a sales workflow with no real mechanism for a human to correct a wrong answer and have that correction stick. Employees notice the tool never actually improves, quietly route around it within a few weeks, and the pilot dies without anyone formally declaring it dead. That's a workflow-integration failure, not a capability failure. It's the same failure mode we cover in how to actually automate a business process with AI agents: an agent has to be wired into the real system of record with a real trigger and a real feedback loop, not dropped in next to the process as a separate chat window that everyone eventually forgets to open.
Buying beats building, by a wide margin

One of the report's more actionable findings is about how companies got their pilots off the ground at all. Pilots built through an external vendor partnership reached real deployment at roughly twice the rate of pilots built entirely in-house, 67% versus 33%, per the report as summarized by Legal.io. That's not an argument for never building anything yourself. It's an argument against making your first AI pilot the place your team learns integration lessons a vendor has already learned on someone else's budget. A vendor that has shipped the same category of workflow a dozen times has already hit the edge cases that kill a pilot quietly in month two. A team building its first agent in-house hits those same edge cases one at a time, on your clock, at your expense.
The 95% versus the 5%, side by side
Pulling the report's findings into one table makes the pattern harder to miss:
| Signal | The 95% (stalled pilots) | The 5% (measurable P&L impact) |
|---|---|---|
| Build approach | Built entirely in-house, first attempt | Bought through an external vendor partnership |
| Deployment rate | 33% reach real deployment | 67% reach real deployment |
| Feedback mechanism | No retained feedback, no adaptation over time | Explicit mechanism for corrections to persist |
| Budget target | Sales and marketing, the visible functions | Back-office operations and finance, the boring ones |
| Review cadence | No hard kill date, becomes a permanent line item | 90-day checkpoint with a real kill criterion |
Every row in that table is a decision made before the model ever generates a single output. That is exactly why "which model should we use" is the wrong first question for a founder evaluating a pilot.
Your budget is probably pointed at the wrong function
The report also found a mismatch between where GenAI budgets go and where the returns actually show up. As Legal.io's summary of the report puts it, enterprise AI budgets "overwhelmingly favor sales and marketing, despite better ROI in operations and finance." Back-office automation, cutting the cost of business-process outsourcing and external agencies, delivered faster payback windows than the customer-facing tools that soak up most of the funding and most of the attention in a board deck. Sales and marketing tools are visible and demo well in a slide. Back-office automation is boring and nobody photographs it, which is probably exactly why it stays underfunded relative to what it actually returns. If you're deciding where to point your first serious AI budget line, the unglamorous back-office workflow is very likely the higher-ROI bet, not the customer-facing one.
A founder's checklist before you sign anything
Before committing real budget to your next AI pilot, run it through five questions:
- Is the target a single, narrow workflow with a metric you can already measure today, before any AI touches it?
- Are you buying from a vendor with a live reference customer doing the same workflow, rather than building the first version yourselves?
- Does the tool have an explicit mechanism for someone to correct a wrong output and have that correction persist, not just a thumbs-down button that goes nowhere?
- Is it wired into the system of record your team already works in every day, instead of living in a separate chat window next to it?
- Have you set a 90-day checkpoint with a real kill criterion, not just a renewal date buried in the invoice?
What this means for your first AI investment
None of this means AI pilots are a bad bet. It means most of them are run like science-fair experiments instead of the same kind of investment a founder would make in any other function: narrow scope, a baseline metric measured before you start, a vendor with proof it has solved your exact problem before, and a hard date to decide whether it's working. The 5% of companies MIT found actually capturing value did not have access to better models than everyone else chasing the same 95% failure rate. They had a better process for deciding what to pilot and how to judge it honestly, which is exactly the discipline we build out in AI Leadership for Founders: how to run a pilot you can actually kill or scale on schedule, instead of one that quietly becomes a permanent line item nobody in the company can explain a year later.
The MIT findings at a glance
| Finding | Number | Source |
|---|---|---|
| GenAI pilots with no measurable P&L impact | 95% | MIT NANDA, "The GenAI Divide" (Jul 2025) |
| Estimated enterprise GenAI spend already committed | $30-40B | Forbes coverage of the report |
| Deployment rate, vendor-partnered pilots | 67% | Legal.io summary of the report |
| Deployment rate, fully in-house pilots | 33% | Legal.io summary, ibid. |
| Orgs actively scaling agentic AI in any function (separate 2025 survey) | 23% | McKinsey, State of AI 2025 |
That last row is worth sitting with: McKinsey's independent, much larger 2025 global survey of nearly 2,000 organizations found almost the same shape of result as MIT's, on a completely separate dataset, using a different methodology. Two research teams measuring different companies arrived at the same conclusion: adoption is nearly universal, scaled value is rare. That's not one report's noise, it's a pattern.
If you're deciding what to automate first so you land in the 5%, not the 95%, the same "repeated, ruled, recoverable" filter from how to actually automate a business process with AI agents is the practical starting point.
Go deeper
AI Leadership for Founders
Want the full AI Leadership for Founders course, not just this post?
Join the waitlist and get a 20% launch discount the moment we open checkout. No payment now.
Taught by Aditya Jha · 40+ AI products shipped for real clients. No spam, unsubscribe any time.
Or join our free community for AI tips while you wait
One email a week: the AI tools, tactics, and course drops actually worth your time. No spam, unsubscribe anytime.
Questions about this or which course fits? Email academy@aibootstrapper.com and we'll answer it directly, not with a support ticket.