Skip to main content

"80% of AI projects fail" traces to a footnote one line long

Published: August 21, 20267 min read
#ai#evidence#research#project-failure
A citation chain traced backwards in four cards. What circulates: RAND found that 80% of AI projects fail. Hop one, the 2024 RAND report, hedged as by some estimates, built on 65 interviews with no denominator. Hop two, footnote 13, pointing to a Fortune vendor profile. Hop three, the bottom: a slew of unnamed surveys measuring business leaders opinions, given as a range of 83 to 92 per cent.
The chain, followed upstream one hop at a time, until it runs out.

You have seen the number. Eighty per cent of AI projects fail, roughly twice the rate of conventional IT projects. It is in the pitch decks, the board papers, the conference keynotes and the LinkedIn posts. I have seen it used to justify buying software and to justify not buying software, which should have been the first clue.

I went looking for the study behind it. There isn't one.

The chain

The citation almost everyone reaches for is RAND. Their 2024 report, The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed, says it twice: "By some estimates, more than 80 percent of AI projects fail."

Read that sentence again, because RAND wrote it carefully. By some estimates. That hedge is theirs, and it is doing a lot of work.

The footnote attached to it is complete. It reads, in full:

  1. Kahn, "Want Your Company's AI Project to Succeed?"

That is Jeremy Kahn, writing in Fortune on 26 July 2022. Not a study. A magazine article, and specifically a profile of the chief executive of a company that sells AI software.

So I read Kahn. Here is the sentence the whole edifice rests on:

That's borne out in a slew of recent surveys, where business leaders have put the failure rate of A.I. projects at between 83% and 92%.

A slew of recent surveys. Unnamed. And look at what they measured: not projects, but business leaders' opinions about projects. Not a rate, but a range, 83 to 92 per cent.

That is the bottom of the chain. There is nothing underneath it.

What happened on the way up

Now watch the number climb back.

Unnamed surveys of executive opinion become "between 83% and 92%" in a vendor profile. RAND rounds that down to "more than 80 percent" and attaches a hedge. And then every downstream citation strips the hedge and reattributes the whole thing, so that what circulates is "RAND found that 80% of AI projects fail".

RAND found no such thing. RAND generated no failure-rate estimate at all, and their study design could not have produced one. It is an exploratory interview study about causes, built on 65 interviews, in which failure is defined perceptually as "a project that was perceived to be a failure by the organization". There is no denominator of projects anywhere in it. You cannot compute a rate without a denominator, and they never claimed to.

RAND's handling of this figure is entirely defensible. What happened to it afterwards is not. The error is not in the research. It is in the transmission, and every time I traced one of these, it had moved in the direction that made the sentence tidier.

The comparison half is worse

"Twice the rate of IT projects" has its own provenance, and it is thinner still.

Two cards side by side. The AI figure: 80%, from 2023, quoted as some estimates place the failure rate as high as 80%, no source given. The IT comparator: not stated, described only as corporate IT project failures from a decade ago, so roughly 2013, no figure and no source given. A dashed line between them is labelled ten years apart. Below, a highlighted panel reads: the clause that never survives the retelling is "from a decade ago", stated plainly by the author and absent from every restatement found.

It traces to Iavor Bojinov in Harvard Business Review, November-December 2023:

Some estimates place the failure rate as high as 80% — almost double the rate of corporate IT project failures from a decade ago.

No source is given for either number. Not for the 80, not for the double.

And notice the clause that never survives the retelling: from a decade ago. Bojinov is comparing a 2023 AI figure against IT failures from around 2013. A ten-year gap, stated plainly by the author, and it vanishes in every restatement I found. The comparison you have heard is between a number with no source and a different number with no source, measured a decade apart.

There is no dataset. There is no consistent definition. There are no two comparable samples. The numerator and the denominator were never measured by the same instrument, or, as far as I can establish, by any instrument at all.

The most repeated empirical claim in this field is not an empirical claim.

What this does and does not show

I want to be careful here, because the tempting conclusion is the wrong one.

This does not prove that AI projects succeed. It does not prove the real rate is low. I have no idea what the real rate is, and after several weeks inside this literature I am fairly confident nobody else does either.

What it proves is narrower and, I think, more useful: the number everyone repeats is not evidence about anything, and nobody checked it for four years.

That last part is the bit I keep returning to. This was not hard. It took me an afternoon, a copy of the RAND PDF and the ability to read a footnote. Anyone quoting the figure could have done the same at any point since 2022. The chain was never hidden. It was simply never followed.

The bit that made me laugh, then stop laughing

While researching this, I used several AI-generated syntheses of the failure literature as a prior-art survey, to see what the consensus looked like before I went at the primary sources.

Two of them independently cited the same study: a RAND review of "over 2,400 enterprise initiatives", complete with a percentage breakdown of causes.

Two cards. Left, highlighted in red: what two AI models reported, 2,400+ enterprise initiatives reviewed by RAND with a percentage split of failure causes, marked fabricated both times with the same figure. Right: what RAND actually did, 65 interviews, reviewing no enterprise initiatives and reporting no cause split, marked verified from the report. Below: this is how a number enters circulation, it arrives sounding citable and everyone downstream assumes somebody upstream checked.

RAND conducted 65 interviews. It reviewed no enterprise initiatives. The study does not exist.

Two separate models produced a citable-sounding artefact with the same fictitious sample size, which is a fairly good illustration of how a number enters circulation and then gets repeated by people who assume somebody upstream checked. The machines are now doing to this literature exactly what the literature has been doing to itself.

What to do with a statistic

I am not going to tell you to be sceptical, which is advice nobody has ever acted on. Here is something narrower you can actually do.

The next time you meet a headline failure statistic, before you repeat it, ask four things. What does it count: projects, companies, use cases, or opinions? On what unit, and is that the same unit as the thing you are comparing it to? Over what window? And is it a measurement of something that happened, or a forecast of something that might?

Four numbered questions to ask before repeating a statistic. One, what does it count: projects, companies, use cases, models or opinions, which are not interchangeable. Two, on what unit, and is it the same unit as whatever you are comparing it to. Three, over what window, and often none is stated at all. Four, highlighted in red, is it a measurement of something that happened or a forecast: of nineteen headline figures traced for the paper, seven were not measurements of anything that happened.

Four questions. They take about a minute, and most numbers do not survive them.

The 80% does not survive the first one. It counts opinions.


This is the first of several pieces drawn from a working paper I published this week, which traces nineteen headline AI failure statistics to their primary sources and classifies each one. Seven of the nineteen turn out not to be measurements of anything that happened. The full paper, with every source stated and every claim checked, is here.

No paywall, no sponsors. If this saved you some time, you can buy me a coffee.

☕ Buy me a coffee

Get the next one in your inbox

I build with AI in the open and write up what held and what didn't. Real numbers, the failures before the wins.

Share this post