What the MIT report on failed AI pilots actually measured

TopicEvidence Published2026-08-23 Read7 min

MIT Project NANDA's State of AI in Business 2025 found that around 5% of custom enterprise AI tools reached production with sustained productivity or P&L impact. That is not the same claim as 95% of GenAI pilots failing, which is how it was reported. The evidence base is 52 executive interviews, 153 survey responses collected at four conferences, and a review of 300 or so publicly disclosed initiatives. The report is not peer reviewed and its method has been publicly challenged. The finding worth your attention is a different one: externally sourced tools reached deployment around 67% of the time against 33% for tools built in house.

In August 2025 a report from MIT knocked value off AI-adjacent share prices and gave every sceptic in every boardroom the same sentence: ninety five per cent of AI pilots fail. It is still being quoted a year later, usually by people selling something.

I read the report. The number is real, the study is real, and the claim being repeated is not the claim the study makes.

What the report is

The GenAI Divide: State of AI in Business 2025, from MIT’s Project NANDA. Published July 2025, twenty six pages, fieldwork carried out between January and June 2025.

The evidence base, as stated in the report:

  • 52 structured interviews with executives
  • 153 survey responses from senior leaders, gathered at four conferences
  • A review of more than 300 publicly disclosed AI initiatives

It has not been peer reviewed.

That is a reasonable amount of work for a six month study. It is also a good deal less than most people picture when they hear “MIT found”. If you have quoted this at a board, it is worth knowing that the survey half of it is 153 people who were standing at a conference.

What it actually found

The finding is that around 5 per cent of custom enterprise AI tools reached production with what the report calls a marked and sustained productivity or profit and loss impact.

Custom enterprise AI tools. Not pilots. Not GenAI programmes. Tools built inside a company for its own use, judged on whether they made it to production and shifted the numbers in a way that stuck.

Flip that to ninety five per cent and broaden the population to every AI pilot everywhere, and you have a much bigger, much scarier claim resting on a much narrower piece of evidence. That is the version that travelled.

The criticism

There is a serious challenge to the method, and anyone using this report should know about it.

The survey responses were collected from executives at conferences. People volunteer different things in that setting than they would in a confidential structured study, and the sample selects for whoever chose to attend and chose to answer. Futuriom published a direct rebuttal in August 2025 arguing the report paints an unfounded picture of enterprise AI. Others have pointed out that six months is a short window in which to judge whether something has produced sustained profit and loss impact, particularly for tools that only went live in the second half of that window.

None of this makes the report worthless. It makes it a piece of evidence with known limits, which is what most business research is. The problem was never the study. It was the reporting.

The finding that deserved the headline

Buried in the same report is a result that is more useful than the one everybody quoted.

Externally sourced tools reached deployment around 67 per cent of the time. Internally built tools reached deployment around 33 per cent of the time.

Roughly twice the success rate for buying rather than building. That is a decision you can act on next week, and it comes with a caveat worth holding onto: the companies doing well with bought tools were not behaving like software shoppers. They treated the vendor relationship the way you would treat an outsourcing partner, demanding customisation to their own processes, holding the supplier to business outcomes rather than model benchmarks, and staying with them through early failures.

For a People team that is a procurement lesson as much as a technology one. The question is not whether to buy or build. It is whether you have the appetite to hold a vendor to an outcome for a year, because the evidence says that is what separates the tools that land from the ones that get quietly renewed and never used.

Why pilots stall, according to the report

The report’s own explanation is worth quoting because it is not the one people assume:

“Tools fail not because of poor models, but because they don’t learn, adapt, or integrate.”

Not model quality. Integration and feedback. A tool that cannot remember what happened last time, cannot take correction, and does not sit inside the workflow stays a productivity toy no matter how good the underlying model is.

The report also notes that more than half of budgets go to visible functions like sales and marketing, despite better returns from back-office automation. For a People function that is close to good news, since most of what you do is exactly the kind of back-office work that the money is not currently chasing.

How to use this properly

If you want to use the 95 per cent figure, use it like this: say what it measured, give the sample size in the same sentence, and say that the method has been disputed. It takes about fifteen seconds and it changes how the room reads you. Anyone who has done their own reading will recognise that you have done yours.

Do not open with it. A pitch that begins with everybody else failing is a pitch that has not yet said anything about the buyer.

And if you take one thing from the report, take the buy-versus-build number, because that one holds up better and it tells you what to do differently on Monday.

Sources

Every figure above traces to one of these. Where a source is contested or its method has been challenged, that is said in the piece rather than left out.

  1. MIT Project NANDA, The GenAI Divide: State of AI in Business 2025 · 2025

    The primary report. 26 pages, fieldwork January to June 2025.

  2. Futuriom, Why We Don't Believe MIT NANDA's Weird AI Study · 2025

    Direct methodological criticism, published August 2025.

  3. An MIT NANDA report misread by all · 2025

    Documents the gap between what the report says and how it was reported.

Questions

Did MIT really find that 95% of AI pilots fail?

Not quite. The report found that around 5% of custom enterprise AI tools reached production with a marked and sustained productivity or P&L impact. The population is custom-built enterprise tools, not GenAI pilots in general, and the difference between those two populations is large.

How big was the sample in the MIT NANDA report?

52 structured executive interviews, 153 survey responses from senior leaders gathered at four conferences, and a review of more than 300 publicly disclosed AI initiatives. Fieldwork ran January to June 2025. The report is not peer reviewed.

Should I stop quoting the 95% figure?

Quote it precisely or not at all. If you use it, say what population it measured and give the sample size in the same breath, and mention that the method has been publicly disputed. Do not open a pitch with it.

What is the most useful finding in the report?

That externally sourced tools reached deployment around 67% of the time compared with around 33% for tools built internally. It is a practical result you can act on, and it received far less coverage than the headline figure.


Elsewhere

The library is the other half of this: small tools built for People teams, published with the repo and the prompts. The newsletter carries one build a week.