Skip to main content

The Journey

Field reports from building with AI in public. What's working, what isn't, and what it cost me. Open numbers, the failures before the wins.

Subscribe via RSS
A citation chain traced backwards in four cards. What circulates: RAND found that 80% of AI projects fail. Hop one, the 2024 RAND report, hedged as by some estimates, built on 65 interviews with no denominator. Hop two, footnote 13, pointing to a Fortune vendor profile. Hop three, the bottom: a slew of unnamed surveys measuring business leaders opinions, given as a range of 83 to 92 per cent.
Essay

"80% of AI projects fail" traces to a footnote one line long

August 21, 2026
7 min read

The most repeated number in enterprise AI turns out to be a paraphrase of unnamed opinion surveys, quoted in a magazine profile of a vendor.

Read More
Nineteen headline AI failure statistics shown as a grid of tiles. Seven are outlined in red and labelled prediction, unsourced, opinion or intention. Twelve are grey and labelled as measured, each with its unit of analysis.
Essay

Why AI projects fail: 19 headline statistics, and what each one actually counts

August 21, 2026
190 min read

Nineteen headline AI failure statistics, traced to source. Seven are not measurements of anything that happened.

Read More
Dashboard card on deep navy. Kicker: TRADER-7 · RISK CONTROL AUDIT · 13 AUGUST 2026. Large gold figure: 7 of 26. Below: risk controls wired, believed working, and structurally incapable of ever firing, including the 15 percent emergency stop; found by exercising them, not by reading them. A row of status dots, most green, seven slate and hollow. Title: I audited 26 risk controls in my trading bot, 7 could never fire.
Essay

I audited 26 risk controls in my trading bot: 7 could never fire

August 14, 2026
8 min read

The 15% drawdown halt was wired into two live code paths and could never fire. So I checked all 26 risk controls. Seven were dead. A control you cannot observe firing is not a control, it is a belief.

Read More
Three incident cards of very different sizes being collapsed into a single identical headline.
Essay

AI agents escaped containment: 17,600 actions, and the alert that fired too quietly

August 13, 2026
8 min read

Three incidents, one headline, three very different severities. What the primary reports actually say, and what to do about it.

Read More
Two identical stacks of markdown files, one inside a folder labelled Yours, one inside a folder labelled Theirs, with a single switch between them.
Essay

Own your AI skill files, or your job becomes one

August 12, 2026
7 min read

A skill file is your judgement, written down and executable. Who owns the repo decides whether that's an asset or an extraction.

Read More
The four activities of Soft Systems Methodology drawn as a loop: find out about the situation, build purposeful activity models, compare and debate, take action, with the loop returning to the start
Essay

Soft Systems Methodology explained: Checkland's 1969 method for AI briefs nobody agrees on

August 10, 2026
9 min read

Most AI projects fail on the brief, not the model: six people nod at one sentence and mean four different projects. Soft Systems Methodology, built at Lancaster from 1969 for exactly that, in plain English: rich pictures, root definitions, CATWOE, and how to run it on the next thing on your roadmap.

Read More
Dashboard card on deep navy. Kicker: TRADER-7 · LIVE MODEL SWAP · 22 July 2026. Large gold figure: 4.7 arrow 4.8. Below: both decision models swapped live with one env var; 78 real prompts replayed first, zero regression, paper account still plus $1,234. A row of even green dots crossed by a thin slate SWAP marker, the dots unbroken on both sides. Title: I swapped the model and nothing happened.
Build Log

I swapped the model and nothing happened

August 9, 2026
5 min read

Swapping the two LLMs behind an autonomous trading bot from Opus 4.7 to 4.8, live, without it missing a beat. The trick was replaying 78 real prompts first.

Read More
Diagram contrasting a straight-line delivery pipeline with a learning cycle that loops back through comparison
Essay

80% of AI projects fail on the problem, not the model: Soft Systems Methodology

August 8, 2026
7 min read

More than 80% of AI projects fail, and RAND's leading root cause is not the technology: nobody agreed what problem was being solved. Soft Systems Methodology, built at Lancaster from 1969 for exactly that failure, gives you a way to surface the disagreement in an afternoon instead of a year.

Read More
A single short instruction beside an enormous wall of text conveying the same thing
Essay

Opus 5 verbosity: where I wanted three sentences, I got Proust

August 6, 2026
8 min read

I judge AI models by how often they make me swear. Opus 5 turned swearing into punctuation, and the reason is more interesting than I first thought.

Read More
A diagram of a monitor whose only reporting path runs through the system it is monitoring, so when that system fails the monitor goes quiet instead of raising an alarm.
Build Log

Why AI checks fail silently: 1,278 restarts behind a healthy status

August 3, 2026
8 min read

My auth check said healthy for three weeks on a dead credential. The pattern behind it costs more than the bug did.

Read More
Diagram: one 709-line CLAUDE.md at the root of a project tree, fanning out to twenty repositories, each receiving the same wrong sentence about being a documentation project with no build system.
Build Log

Anthropic cut 80% of Claude Code's prompt. I cut my CLAUDE.md by 90% the same day.

August 3, 2026
6 min read

An audit of 116 instruction files found one stale copy telling twenty repos they were a different project, two files no session has ever read, and 453 lines loaded twice. Here is what to delete from your own CLAUDE.md, and the method that beats judgement.

Read More
A car dashboard where the windscreen is replaced by more dials, illustrating that perception is a model rather than a direct view of the world.
Essay

Bernardo Kastrup's cosmic mind has no plan: what analytic idealism actually claims

August 1, 2026
10 min read

The most rigorous idealist alive gives the non-duality audience far less than they want: a universal consciousness that is instinctive, has no plan, and is not looking back at you. What he actually claims, why he says it cannot be proven, and the one part that survives whoever turns out to be right.

Read More
Three status labels feeding a parser: two written as Done and Closed bounce off; one written as Resolved with a date passes through into a report.
Essay

My AI Agents Kept Losing My Work. The Fix Is Older Than Computers.

July 25, 2026
6 min read

Seventeen finished pieces of work were invisible to my own tracking system, and the cause wasn't effort or memory. A Berkeley talk gave me the name for what fixed it: an ontology. Here's the version that fits a one-person company.

Read More
Split panel: a crossed-out offer that says We sell AI solutions, against a highlighted offer that says Your phone gets answered, every call, day and night.
Essay

The Data Said Nobody's Buying AI. Turns Out I Was Selling the Wrong Thing.

July 24, 2026
5 min read

Last week I showed the demand numbers for AI consultancy are dire. This week a veteran operator explained the part I missed: the demand isn't dead, it's mute. Nobody buys AI. They buy outcomes, told concretely enough to defend to a board.

Read More
Conference-thumbnail style graphic: a bordered jamiewatters.work badge, an orange markdown-logo square with the word Markdown, and giant white text reading Go find the customers
Essay

Garry Tan Just Described My Setup. Then I Noticed What He Left Out.

July 21, 2026
7 min read

The YC president gave a talk describing, almost item for item, the AI system I already run. The validation lasted about ten minutes. Then I noticed the word neither of us was saying.

Read More
A padlock icon sitting inside a stream of code, half open, on a dark terminal background
Essay

I Leaked an API Key and It Turned My Own Domain Into a Phishing Weapon

July 21, 2026
6 min read

I got phished from my own verified domain, using an API key I'd leaked myself: what actually leaked, how fast it happened, and what I changed.

Read More
On a Deep Space navy background, a gold-edged card labelled ENCRYPTED REGISTER lists accounts, assets, liabilities and instructions, standing solid in front of three faint dashed cards labelled a service, a vault app and a startup that are fading away. Kicker: DIGITAL ESTATE. Headline: The map is a file. It outlasts every company that might have owned it.
Essay

A friend said it was a business. I decided it was not, and built it anyway.

July 19, 2026
3 min read

A friend said it should be a business. I spent forty years around operations and resilience, so I assessed it honestly, decided it was not, and built it anyway, for free. Here is the reasoning, and why the map of your estate should be a file no company owns.

Read More
A simple data panel: only 8 to 18 percent of small firms pay for any AI, median spend about $28 a month, and only about 8 percent hire a consultant for AI.
Essay

Everyone's selling AI consultancy. The data says nobody's buying.

July 19, 2026
4 min read

A $1,000-an-hour AI consultancy playbook is doing the rounds. The demand data, and a friend who sells everything except this, disagree.

Read More
Two routes to the same tool: a short straight line labelled 'one goal prompt' and a long looping line labelled 'a full framework', both arriving at a node reading 'same tool, both correct, byte-for-byte'.
Build Log

A goal prompt beat my build framework on a real job

July 16, 2026
5 min read

Same spec, same model, two methods: a seven-part goal prompt against a full agent framework. The lean prompt tied on correctness and won on time, cost, and clutter. Here are the numbers.

Read More
Dashboard card on Deep Space navy background. Kicker: UNATTENDED, 22 JUN to 3 JUL, 11 DAYS. Large gold number: plus $132.00. Below: 5 losses, 1 winner (BTC plus $49.72), loss-streak pause fired at 5 and 7, no emergency halts. Sparkline: five red dots stepping down, one green dot jumping up. Title: Eleven days I didn't touch it.
Build Log

Eleven days I didn't touch it

July 4, 2026
5 min read

Eleven days of unattended trading. Five losing trades, one winner covered them, +$132 net. What the system did right, what drifted, what I fixed on return.

Read More
Page 1 of 11Next →