ACADEMY
All posts

Why 95% of AI Pilots Fail to Show ROI (And How to Be in the 5%)

MIT's 2025 study of 300 enterprise AI deployments found 95% never move the P&L. Here's what the other 5% actually did differently, and what it means before you sign your next AI vendor contract.

In July 2025, MIT's Project NANDA published a study that should worry anyone about to sign an AI vendor contract. After reviewing 300 public AI deployments, surveying 153 business leaders, and interviewing 52 executives, the researchers found that 95% of generative AI pilots inside companies were producing no measurable effect on profit or loss, despite an estimated $30-40 billion in enterprise spending already committed. This wasn't a study about whether AI works. It was a study about why most companies fail to make it work, and the answer has almost nothing to do with which model they picked.

What MIT actually measured

Researchers at MIT's NANDA initiative didn't just poll opinions. Across 300 public deployments and 52 executive interviews, they separated "adoption," a team tried a tool, from "transformation," the tool changed a real P&L line. The gap between those two numbers is what the report, titled "The GenAI Divide: State of AI in Business 2025," calls out directly: nearly every mid-size and large company has piloted something by now, but as Forbes reported on the study, only 5% of those pilots ever reach the point where a finance team can point to a number on a statement and say "that came from the AI."

The real problem isn't the model, it's the learning gap

The instinct when a pilot underperforms is to blame the model, wait for the next release, and try again. MIT's researchers found that's the wrong diagnosis. Their core finding was that most of the GenAI systems that failed to move the needle "do not retain feedback, adapt to context, or improve over time." In practice, that means a chatbot gets bolted onto a support queue or a sales workflow with no real mechanism for a human to correct a wrong answer and have that correction stick. Employees notice the tool never actually improves, quietly route around it within a few weeks, and the pilot dies without anyone formally declaring it dead. That's a workflow-integration failure, not a capability failure. It's the same failure mode we cover in how to actually automate a business process with AI agents: an agent has to be wired into the real system of record with a real trigger and a real feedback loop, not dropped in next to the process as a separate chat window that everyone eventually forgets to open.

Buying beats building, by a wide margin

Source: MIT NANDA, "The GenAI Divide: State of AI in Business 2025" (July 2025), as reported by Forbes and Legal.io, August 2025.
Source: MIT NANDA, "The GenAI Divide: State of AI in Business 2025" (July 2025), as reported by Forbes and Legal.io, August 2025.

One of the report's more actionable findings is about how companies got their pilots off the ground at all. Pilots built through an external vendor partnership reached real deployment at roughly twice the rate of pilots built entirely in-house, 67% versus 33%, per the report as summarized by Legal.io. That's not an argument for never building anything yourself. It's an argument against making your first AI pilot the place your team learns integration lessons a vendor has already learned on someone else's budget. A vendor that has shipped the same category of workflow a dozen times has already hit the edge cases that kill a pilot quietly in month two. A team building its first agent in-house hits those same edge cases one at a time, on your clock, at your expense.

The 95% versus the 5%, side by side

Pulling the report's findings into one table makes the pattern harder to miss:

SignalThe 95% (stalled pilots)The 5% (measurable P&L impact)
Build approachBuilt entirely in-house, first attemptBought through an external vendor partnership
Deployment rate33% reach real deployment67% reach real deployment
Feedback mechanismNo retained feedback, no adaptation over timeExplicit mechanism for corrections to persist
Budget targetSales and marketing, the visible functionsBack-office operations and finance, the boring ones
Review cadenceNo hard kill date, becomes a permanent line item90-day checkpoint with a real kill criterion

Every row in that table is a decision made before the model ever generates a single output. That is exactly why "which model should we use" is the wrong first question for a founder evaluating a pilot.

Your budget is probably pointed at the wrong function

The report also found a mismatch between where GenAI budgets go and where the returns actually show up. As Legal.io's summary of the report puts it, enterprise AI budgets "overwhelmingly favor sales and marketing, despite better ROI in operations and finance." Back-office automation, cutting the cost of business-process outsourcing and external agencies, delivered faster payback windows than the customer-facing tools that soak up most of the funding and most of the attention in a board deck. Sales and marketing tools are visible and demo well in a slide. Back-office automation is boring and nobody photographs it, which is probably exactly why it stays underfunded relative to what it actually returns. If you're deciding where to point your first serious AI budget line, the unglamorous back-office workflow is very likely the higher-ROI bet, not the customer-facing one.

A founder's checklist before you sign anything

Before committing real budget to your next AI pilot, run it through five questions:

  • Is the target a single, narrow workflow with a metric you can already measure today, before any AI touches it?
  • Are you buying from a vendor with a live reference customer doing the same workflow, rather than building the first version yourselves?
  • Does the tool have an explicit mechanism for someone to correct a wrong output and have that correction persist, not just a thumbs-down button that goes nowhere?
  • Is it wired into the system of record your team already works in every day, instead of living in a separate chat window next to it?
  • Have you set a 90-day checkpoint with a real kill criterion, not just a renewal date buried in the invoice?

What this means for your first AI investment

None of this means AI pilots are a bad bet. It means most of them are run like science-fair experiments instead of the same kind of investment a founder would make in any other function: narrow scope, a baseline metric measured before you start, a vendor with proof it has solved your exact problem before, and a hard date to decide whether it's working. The 5% of companies MIT found actually capturing value did not have access to better models than everyone else chasing the same 95% failure rate. They had a better process for deciding what to pilot and how to judge it honestly, which is exactly the discipline we build out in AI Leadership for Founders: how to run a pilot you can actually kill or scale on schedule, instead of one that quietly becomes a permanent line item nobody in the company can explain a year later.

The MIT findings at a glance

FindingNumberSource
GenAI pilots with no measurable P&L impact95%MIT NANDA, "The GenAI Divide" (Jul 2025)
Estimated enterprise GenAI spend already committed$30-40BForbes coverage of the report
Deployment rate, vendor-partnered pilots67%Legal.io summary of the report
Deployment rate, fully in-house pilots33%Legal.io summary, ibid.
Orgs actively scaling agentic AI in any function (separate 2025 survey)23%McKinsey, State of AI 2025

That last row is worth sitting with: McKinsey's independent, much larger 2025 global survey of nearly 2,000 organizations found almost the same shape of result as MIT's, on a completely separate dataset, using a different methodology. Two research teams measuring different companies arrived at the same conclusion: adoption is nearly universal, scaled value is rare. That's not one report's noise, it's a pattern.

If you're deciding what to automate first so you land in the 5%, not the 95%, the same "repeated, ruled, recoverable" filter from how to actually automate a business process with AI agents is the practical starting point.

Go deeper

AI Leadership for Founders

Courses launching soon

Want the full AI Leadership for Founders course, not just this post?

Join the waitlist and get a 20% launch discount the moment we open checkout. No payment now.

Taught by Aditya Jha · 40+ AI products shipped for real clients. No spam, unsubscribe any time.

Or join our free community for AI tips while you wait

AI Weekly Radar

One email a week: the AI tools, tactics, and course drops actually worth your time. No spam, unsubscribe anytime.

Questions about this or which course fits? Email academy@aibootstrapper.com and we'll answer it directly, not with a support ticket.