"80% of AI projects fail" traces to a footnote one line long

You have seen the number. Eighty per cent of AI projects fail, roughly twice the rate of conventional IT projects. It is in the pitch decks, the board papers, the conference keynotes and the LinkedIn posts. I have seen it used to justify buying software and to justify not buying software, which should have been the first clue.
I went looking for the study behind it. There isn't one.
The chain
The citation almost everyone reaches for is RAND. Their 2024 report, The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed, says it twice: "By some estimates, more than 80 percent of AI projects fail."
Read that sentence again, because RAND wrote it carefully. By some estimates. That hedge is theirs, and it is doing a lot of work.
The footnote attached to it is complete. It reads, in full:
- Kahn, "Want Your Company's AI Project to Succeed?"
That is Jeremy Kahn, writing in Fortune on 26 July 2022. Not a study. A magazine article, and specifically a profile of the chief executive of a company that sells AI software.
So I read Kahn. Here is the sentence the whole edifice rests on:
That's borne out in a slew of recent surveys, where business leaders have put the failure rate of A.I. projects at between 83% and 92%.
A slew of recent surveys. Unnamed. And look at what they measured: not projects, but business leaders' opinions about projects. Not a rate, but a range, 83 to 92 per cent.
That is the bottom of the chain. There is nothing underneath it.
What happened on the way up
Now watch the number climb back.
Unnamed surveys of executive opinion become "between 83% and 92%" in a vendor profile. RAND rounds that down to "more than 80 percent" and attaches a hedge. And then every downstream citation strips the hedge and reattributes the whole thing, so that what circulates is "RAND found that 80% of AI projects fail".
RAND found no such thing. RAND generated no failure-rate estimate at all, and their study design could not have produced one. It is an exploratory interview study about causes, built on 65 interviews, in which failure is defined perceptually as "a project that was perceived to be a failure by the organization". There is no denominator of projects anywhere in it. You cannot compute a rate without a denominator, and they never claimed to.
RAND's handling of this figure is entirely defensible. What happened to it afterwards is not. The error is not in the research. It is in the transmission, and every time I traced one of these, it had moved in the direction that made the sentence tidier.
The comparison half is worse
"Twice the rate of IT projects" has its own provenance, and it is thinner still.

It traces to Iavor Bojinov in Harvard Business Review, November-December 2023:
Some estimates place the failure rate as high as 80% — almost double the rate of corporate IT project failures from a decade ago.
No source is given for either number. Not for the 80, not for the double.
And notice the clause that never survives the retelling: from a decade ago. Bojinov is comparing a 2023 AI figure against IT failures from around 2013. A ten-year gap, stated plainly by the author, and it vanishes in every restatement I found. The comparison you have heard is between a number with no source and a different number with no source, measured a decade apart.
There is no dataset. There is no consistent definition. There are no two comparable samples. The numerator and the denominator were never measured by the same instrument, or, as far as I can establish, by any instrument at all.
The most repeated empirical claim in this field is not an empirical claim.
What this does and does not show
I want to be careful here, because the tempting conclusion is the wrong one.
This does not prove that AI projects succeed. It does not prove the real rate is low. I have no idea what the real rate is, and after several weeks inside this literature I am fairly confident nobody else does either.
What it proves is narrower and, I think, more useful: the number everyone repeats is not evidence about anything, and nobody checked it for four years.
That last part is the bit I keep returning to. This was not hard. It took me an afternoon, a copy of the RAND PDF and the ability to read a footnote. Anyone quoting the figure could have done the same at any point since 2022. The chain was never hidden. It was simply never followed.
The bit that made me laugh, then stop laughing
While researching this, I used several AI-generated syntheses of the failure literature as a prior-art survey, to see what the consensus looked like before I went at the primary sources.
Two of them independently cited the same study: a RAND review of "over 2,400 enterprise initiatives", complete with a percentage breakdown of causes.

RAND conducted 65 interviews. It reviewed no enterprise initiatives. The study does not exist.
Two separate models produced a citable-sounding artefact with the same fictitious sample size, which is a fairly good illustration of how a number enters circulation and then gets repeated by people who assume somebody upstream checked. The machines are now doing to this literature exactly what the literature has been doing to itself.
What to do with a statistic
I am not going to tell you to be sceptical, which is advice nobody has ever acted on. Here is something narrower you can actually do.
The next time you meet a headline failure statistic, before you repeat it, ask four things. What does it count: projects, companies, use cases, or opinions? On what unit, and is that the same unit as the thing you are comparing it to? Over what window? And is it a measurement of something that happened, or a forecast of something that might?

Four questions. They take about a minute, and most numbers do not survive them.
The 80% does not survive the first one. It counts opinions.
This is the first of several pieces drawn from a working paper I published this week, which traces nineteen headline AI failure statistics to their primary sources and classifies each one. Seven of the nineteen turn out not to be measurements of anything that happened. The full paper, with every source stated and every claim checked, is here.