ACADEMY
All posts

Why 95% of AI Pilots Fail to Show ROI (And How to Be in the 5%)

MIT's 2025 study of 300 enterprise AI deployments found 95% never move the P&L. Here's what the other 5% actually did differently, and what it means before you sign your next AI vendor contract.

In July 2025, MIT's Project NANDA published a study that should worry anyone about to sign an AI vendor contract. After reviewing 300 public AI deployments, surveying 153 business leaders, and interviewing 52 executives, the researchers found that 95% of generative AI pilots inside companies were producing no measurable effect on profit or loss, despite an estimated $30-40 billion in enterprise spending already committed. This wasn't a study about whether AI works. It was a study about why most companies fail to make it work, and the answer has almost nothing to do with which model they picked.

What MIT actually measured

Researchers at MIT's NANDA initiative didn't just poll opinions. Across 300 public deployments and 52 executive interviews, they separated "adoption," a team tried a tool, from "transformation," the tool changed a real P&L line. The gap between those two numbers is what the report, titled "The GenAI Divide: State of AI in Business 2025," calls out directly: nearly every mid-size and large company has piloted something by now, but as Forbes reported on the study, only 5% of those pilots ever reach the point where a finance team can point to a number on a statement and say "that came from the AI."

The real problem isn't the model, it's the learning gap

The instinct when a pilot underperforms is to blame the model, wait for the next release, and try again. MIT's researchers found that's the wrong diagnosis. Their core finding was that most of the GenAI systems that failed to move the needle "do not retain feedback, adapt to context, or improve over time." In practice, that means a chatbot gets bolted onto a support queue or a sales workflow with no real mechanism for a human to correct a wrong answer and have that correction stick. Employees notice the tool never actually improves, quietly route around it within a few weeks, and the pilot dies without anyone formally declaring it dead. That's a workflow-integration failure, not a capability failure. It's the same failure mode we cover in how to actually automate a business process with AI agents: an agent has to be wired into the real system of record with a real trigger and a real feedback loop, not dropped in next to the process as a separate chat window that everyone eventually forgets to open.

Buying beats building, by a wide margin

Source: MIT NANDA, "The GenAI Divide: State of AI in Business 2025" (July 2025), as reported by Forbes and Legal.io, August 2025.
Source: MIT NANDA, "The GenAI Divide: State of AI in Business 2025" (July 2025), as reported by Forbes and Legal.io, August 2025.

One of the report's more actionable findings is about how companies got their pilots off the ground at all. Pilots built through an external vendor partnership reached real deployment at roughly twice the rate of pilots built entirely in-house, 67% versus 33%, per the report as summarized by Legal.io. That's not an argument for never building anything yourself. It's an argument against making your first AI pilot the place your team learns integration lessons a vendor has already learned on someone else's budget. A vendor that has shipped the same category of workflow a dozen times has already hit the edge cases that kill a pilot quietly in month two. A team building its first agent in-house hits those same edge cases one at a time, on your clock, at your expense.

Your budget is probably pointed at the wrong function

The report also found a mismatch between where GenAI budgets go and where the returns actually show up. As Legal.io's summary of the report puts it, enterprise AI budgets "overwhelmingly favor sales and marketing, despite better ROI in operations and finance." Back-office automation, cutting the cost of business-process outsourcing and external agencies, delivered faster payback windows than the customer-facing tools that soak up most of the funding and most of the attention in a board deck. Sales and marketing tools are visible and demo well in a slide. Back-office automation is boring and nobody photographs it, which is probably exactly why it stays underfunded relative to what it actually returns. If you're deciding where to point your first serious AI budget line, the unglamorous back-office workflow is very likely the higher-ROI bet, not the customer-facing one.

A founder's checklist before you sign anything

Before committing real budget to your next AI pilot, run it through five questions:

  • Is the target a single, narrow workflow with a metric you can already measure today, before any AI touches it?
  • Are you buying from a vendor with a live reference customer doing the same workflow, rather than building the first version yourselves?
  • Does the tool have an explicit mechanism for someone to correct a wrong output and have that correction persist, not just a thumbs-down button that goes nowhere?
  • Is it wired into the system of record your team already works in every day, instead of living in a separate chat window next to it?
  • Have you set a 90-day checkpoint with a real kill criterion, not just a renewal date buried in the invoice?

What this means for your first AI investment

None of this means AI pilots are a bad bet. It means most of them are run like science-fair experiments instead of the same kind of investment a founder would make in any other function: narrow scope, a baseline metric measured before you start, a vendor with proof it has solved your exact problem before, and a hard date to decide whether it's working. The 5% of companies MIT found actually capturing value did not have access to better models than everyone else chasing the same 95% failure rate. They had a better process for deciding what to pilot and how to judge it honestly, which is exactly the discipline we build out in AI Leadership for Founders: how to run a pilot you can actually kill or scale on schedule, instead of one that quietly becomes a permanent line item nobody in the company can explain a year later.

Go deeper

AI Leadership for Founders