AI vs IT project failure rate: the comparison has never been made
Jamie Watters
Operational resilience and AI delivery practitioner. Technology since 1985.

The claim comes up in most writing about AI projects: they fail at about twice the rate of ordinary IT projects. Traced back, it is one sentence in Harvard Business Review in late 2023, which puts the failure rate "as high as 80%" and calls that "almost double the rate of corporate IT project failures from a decade ago". No source is given for either number. RAND's 2024 report repeats the 80% with a hedge of its own, and everyone downstream repeats both. I went looking for the study behind the comparison while tracing nineteen such statistics to their sources. There isn't one.
Not tested and found true. Not tested and found false. Not tested, as far as a search that chains outward from the field's own most-cited documents can tell, and I will come back to that boundary.
Here is what the comparison would need. Two populations, one AI and one not, measured on one outcome definition, by one instrument, over comparable windows, with a stated denominator on both sides. Nothing like that exists. What exists is an AI figure that paraphrases unnamed opinion surveys, and an IT figure that its own field spent twenty years taking apart.
This piece is about the IT side of that comparison, because it is worse than the AI side, and because the critique already written against it is, line for line, the critique the AI numbers deserve.
The baseline moved 88 points on a definition
Nearly every "most IT projects fail" claim traces to the Standish Group's 1994 CHAOS Report. Its definitions are worth reading slowly. A success was a project "completed on-time and on-budget, with all features and functions as initially specified". That was 16.2% of the sample. A challenged project was "completed and operational but over-budget, over the time estimate, and offers fewer features". Impaired meant cancelled.
Notice what success means there. It means the estimate was accurate. A project that delivered enormous value late is a failure. A project that delivered nothing anyone used, on time and on budget, exactly as specified, is a success.
In 2010, Eveleens and Verhoef published the case against those definitions in IEEE Software, and they did not hedge: the definitions "are misleading, one-sided, pervert the estimation practice, and result in meaningless figures". Then they ran the experiment that should have ended the argument. They applied the Standish definition of success to 5,457 real forecasts across 1,211 projects in three organisations. One of them, Landmark Graphics, habitually underestimated: on its 121 projects the Standish success rate came out at 5.8%. So they built a mirror image of Landmark, every forecast off by the same amount in the opposite direction, overestimates instead of underestimates, and scored it on the same definition. It came out at 94.2%.
Same projects. Same size of forecasting error. Reverse its direction and the headline figure moves by 88 percentage points, because the definition only punishes error in one direction.

It gets worse for the baseline. Four years earlier, Jørgensen and Moløkken-Østvold had quoted Standish's own account of how the 1994 data was collected: surveys sent to IT executives "asking them to share failure stories". A sample solicited as failure stories cannot support a population failure rate. They also found the famous 189% average cost overrun inconsistently defined across Standish's own documents and an outlier against every comparable study of the period, which clustered around 33%. When they asked Standish for its method, the answer was that providing it "would be like giving away their business for free".
Standish did revise the definition in 2015, adding value to it, and the series changed again. So the IT folklore number was solicited as failure stories, defined as estimate accuracy rather than value, redefined mid-series, and moves by 88 points when the direction of the estimating error flips. That is not a baseline. That is a definition wearing a percentage. Whether the HBR sentence had this number or some other rule of thumb in mind is unknowable, because it names none, which is the point.
The mirror
Every defect in CHAOS reappears in the AI failure literature, in the same order, twenty to thirty years later. Failure-solicited sampling becomes conference-recruited practitioner panels. Success as estimate accuracy inside an arbitrary window becomes "no measurable P&L impact within six months". Undisclosed method becomes the proprietary consultancy report and the prediction whose only disclosed input is a webinar poll. Definitions that change between editions while the numbers are presented as a trend become 80%, 85% and 95% lined up as if they measured the same thing.
I traced nineteen of those AI figures to their sources in the review paper this series comes from. Seven of them are not measurements of anything that happened. None of them has had the Eveleens and Verhoef experiment run on it: take a real portfolio, hold its performance constant, vary the success definition, and report how far the failure rate moves. It needs no new fieldwork. Nobody has done it.
What a real comparison looks like, and what it says about IT
One piece of work does properly what the AI literature has not attempted. Flyvbjerg, Budzier, Aaen, Keil and Zottoli compared 23 project types in Project Management Journal, published online in July 2025: 11,011 projects, 126 countries, one outcome measure (cost overrun as the ratio of actual to estimated cost, not feature conformance), and a stated null hypothesis that IT projects are no different from other projects in cost risk. The null was rejected. But where IT ranks is not where the folklore puts it.
On mean cost overrun, IT is fifth of 23, at 1.73. Nuclear storage, the Olympics, nuclear power and hydroelectric dams all sit above it. On the average, IT is bad but unexceptional.
On the tail, IT is the worst of all 23. Its fitted tail index, the Pareto alpha, is 0.92 on the point estimate, the only project type at or below 1, and a tail that heavy has, in plain terms, no upper bound you can plan against: the average and the spread are not stable numbers. Just over 18% of IT projects cost more than 1.5 times their estimate, an overrun above 50%, and the mean overrun ratio inside that group is 5.53, which is 453% over.
An earlier dataset from the same research programme, 5,392 IT projects completed between 2002 and 2014, gives the shape in two numbers. Mean cost-overrun ratio: 1.8. Median: 1.0. The mean differs slightly from the 1.73 above because the two samples are not the same set of projects. The median IT project comes in on budget. The average is produced by a tail, not by the typical project.

So what distinguishes IT is not that it fails more often, or overruns more on average. It is unexceptional on the average and unique in the tail, and the average is made by the tail. When it goes wrong it goes wrong without an upper bound.
Now apply that to AI. AI projects could differ from IT projects in the rate, in the typical outcome, or in the tail. But the IT evidence says the tail is where IT's own distinctiveness lives, so any comparison that reports a rate and nothing else has not looked in the one place a real difference is most likely to hide.
None of the nineteen AI figures I traced comes from a study that reports a distributional shape: no median beside the mean, no tail index, no count of extreme outcomes. That is the corpus reachable by chaining from the field's most-cited documents, not the whole literature, and I would be glad to be sent the paper that does it. Within that corpus the genre reports rates, and a rate cannot see a tail.
The limitation, stated
The best IT baseline in existence predates the phenomenon being compared to it. The 5,392-project dataset runs to 2014. The 23-type dataset contains no meaningful population of generative AI projects and could not, given when its projects completed. When I say IT's distinctiveness is its tail, I am describing projects mostly finished before the transformer architecture existed.
That cuts both ways and I am not going to pretend otherwise. The comparison cannot be made properly today even by someone doing it correctly. It also means anyone asserting AI is worse than IT is asserting it against a baseline they have not read, which predates their subject, and which its own literature contests.
There is one quantified estimate of how much of AI project difficulty is new, and I offer it with its status attached: Bérubé, Giannelia and Vial's Delphi panel of 18 experts found that six of sixteen barriers, 37% as the authors give it, did not appear in earlier IT implementation studies. The panel's ranking did not reach consensus, and of the four barriers with low agreement, the authors note only two are AI-specific. A straw in the wind.
And the search itself has a boundary. The nineteen figures and the studies behind them were found by following citations outward from the documents the field cites most. A comparison published outside that web, in another language or never cited, would not have been found. So the defensible claim is narrower than the title: nothing in the material the field actually circulates makes the comparison, and the material it circulates is what lands on your desk.
What to ask instead
Two questions, whenever "AI projects fail twice as often as IT" reaches your desk.
Which definition? If the answer is estimate conformance inside a fixed window, you are looking at the CHAOS defect with a new label, and the Eveleens and Verhoef result says the number could be almost anything.
Where is the median? A rate tells you nothing about the shape. The one thing IT's own data proves is that the typical project and the average project are different animals, and that the average is made by a tail. Anyone who wants to show AI is worse has to show the shape, not just the rate. Nobody in the circulating literature has, because nobody has measured it.
The claim has been repeated since late 2023. The comparison it rests on has never been made in the material that repeats it. Those are two different facts, and only one of them is about AI.