Skip to main content

80% of AI projects fail on the problem, not the model: Soft Systems Methodology

Published: August 8, 20267 min read
#ssm#checkland#ai-deployment#systems-thinking#project-failure
Diagram contrasting a straight-line delivery pipeline with a learning cycle that loops back through comparison

Four research groups have now measured AI project failure and landed in roughly the same place, and almost everyone has drawn the wrong conclusion from it.

RAND put it at more than 80% of AI projects failing, twice the failure rate of IT projects with no AI in them. That was August 2024, built on 65 interviews with data scientists and engineers. In March 2025, S&P Global Market Intelligence found 42% of businesses had abandoned most of their AI initiatives, up from 17% the year before, across more than 1,000 respondents in North America and Europe. In August 2025, MIT's NANDA initiative reported that around 95% of enterprise generative AI pilots produced no measurable effect on the profit and loss account. Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027.

The popular reading is that the technology is not ready yet. Read the root causes instead.

RAND's leading cause, across every organisation they studied, is misunderstanding or miscommunicating the problem the project was meant to solve. MIT's is that the tools do not learn the way the work is actually done. Gartner's three cancellation reasons are escalating costs, unclear business value, and inadequate risk controls. Only one of those is anywhere near a model.

Take the AI off and you are looking at the oldest failure in the business: an organisation that has agreed on a solution without agreeing on a problem. The models are fine. The briefs are rubbish.


I started as a systems programmer in 1985, writing OS-level code and assembler, and then spent years in project and programme management before landing in business continuity. So I have sat on both sides of this. I have written the code and I have written the plan, and I can tell you that the plan is where the money dies.

The reason is structural. Every delivery method most organisations own assumes the objective is already settled. You gather requirements, you build a proof of concept, you agree success criteria, you roll out. Each of those steps takes the problem as given and optimises the route to it. None of them has a stage whose job is to ask whether the six people in the room mean the same thing by the sentence they all just nodded at.

That gap was diagnosed and given a method in 1969. We have simply not been using it.


Peter Checkland was at Lancaster University's Systems Department, running what became a ten-year action research programme into why systems engineering kept failing. Systems engineering worked beautifully on defined problems: build a plant that produces this output at this cost. Turned loose on management problems it fell apart, because in a management problem the objective is the thing in dispute. There is no agreed goal to engineer towards.

His response was Soft Systems Methodology. The first formal account is a 1972 paper, "Towards a systems-based methodology for real-world problem solving", and the method was worked out in the field until around 1990. Systems Thinking, Systems Practice landed in 1981, Soft Systems Methodology in Action with Jim Scholes in 1990. By 2000 John Mingers could write that SSM had "reoriented an entire discipline and touched the lives of literally thousands of people". It is taught, it is respected, and outside operational research and health informatics almost nobody in technology delivery has heard of it.

The move at the centre of it is small and slightly devious. Checkland stopped treating the organisation as a system and started treating "system" as a mental construct you build in order to learn about a mess. You do not model the company. You model several purposeful activity systems that different people believe the company should be, and then you compare those models against what is actually happening. The comparison is the point. It is a learning cycle, not a solution generator.

The machinery is unglamorous. You draw a rich picture: an ugly hand-drawn sketch of the situation, people, conflicts, flows of work, arrows, speech bubbles, no formal notation, because formal notation makes you leave out the awkward parts. You write root definitions in the form "a system to do X by means of Y in order to achieve Z". You test each one with CATWOE, formalised by Smyth and Checkland in 1976: customers, actors, transformation, Weltanschauung, owner, environmental constraints. The Weltanschauung is the load-bearing one. It is the worldview that makes the transformation meaningful to whoever proposed it, and it is almost always unstated.

Then you build a conceptual model of each root definition, compare it with stage two reality, and argue your way to changes that are both systemically desirable and culturally feasible. That last pair of words has killed more bad projects than any risk register I have ever seen.

Anatomy of a root definition: the X, Y, Z sentence above the six CATWOE elements, with Weltanschauung marked as the element usually left unstated


Here is what it does to an AI brief.

Take a sentence you have heard this year: "deploy an AI agent in customer service". Everyone nods. Now write it as four root definitions, one per worldview in the room.

Operations wants a system to deflect contacts from human agents by means of automated first-line response in order to cut cost per ticket. Customer experience wants a system to shorten time to resolution by means of instant triage in order to protect retention. Compliance wants a system to constrain what is said to customers by means of a bounded response set in order to keep advice inside regulatory limits. And the agents on the floor are living inside a fifth one nobody wrote down: a system to remove the easy tickets by means of automation, in order to leave them a queue of nothing but hard, angry calls, measured against the same average handling time as before.

One brief, four root definitions: operations, customer experience and compliance each stated, plus an unwritten fourth held by the agents on the floor

Those are not one project. They are four, and the fourth one explains the rollout that quietly stalls at 30% adoption while the dashboard says the model is performing well. You can build any of the first three. You cannot build all three at once with the same bounded response set, and the moment you notice that, you are having the argument that would otherwise have surfaced eleven months and a seven-figure budget later.

That is the whole bridge. AI deployment fails at the point where purpose is contested, and SSM is a method built for nothing else.


Two honest qualifications, because this is the part where a piece like this normally starts selling.

SSM stayed niche for a practical reason: it was expensive. Rich pictures, root definitions and conceptual models took days of facilitated workshops, and the output was a better conversation rather than a document you could hand to a delivery team. That cost is now mostly gone. Drafting six root definitions from six worldviews, and stress-testing each with CATWOE, is exactly the sort of structured generation a language model does well in a few minutes. The argument still needs the humans in the room. The paperwork that used to make the argument unaffordable does not.

And SSM will not save you from power. It surfaces disagreement, it does not resolve it. If the sponsor has already decided, no amount of CATWOE outranks them, and the standard academic objection to SSM is precisely this: it assumes participants can debate on something like equal terms, which in most organisations is a fiction. I should also be clear that nobody has run the study that would settle this. There is no published trial showing SSM improves AI project outcomes. What exists is a well-documented failure mode and a fifty-year-old method aimed directly at it. I am arguing from that fit, not from evidence I do not have.

[VERIFY: Jamie, if you want a first-hand example here from your own programme management or continuity work, give me the situation and I will write it in. I have not invented one.]


Before the next AI project starts, do one thing. Ask everyone who will be in the steering group to write the transformation on their own, in one line, in the form "a system to do X by means of Y in order to achieve Z". Do not discuss it first. Collect the lines and read them out.

If two of those lines describe different projects, you have found the failure while it still costs an afternoon.

No paywall, no sponsors. If this saved you some time, you can buy me a coffee.

☕ Buy me a coffee

Get the next one in your inbox

I build with AI in the open and write up what held and what didn't. Real numbers, the failures before the wins.

Share this post