Skip to main content

AI

Working with AI, building with it, and thinking about where it goes.

A decision path diagram showing an escalation route that terminates at the same executive who owns the shipping decision, labelled disclosure rather than control.
Essay

Anthropic's safety policy has never stopped a release. Yours probably hasn't either.

August 31, 2026
12 min read

The best-documented AI safety regime in existence publishes its own failures, raises its own risk grades, and has never stopped a release. Here is the test.

Read More
Two columns. On the left, seven remediation actions each marked closed. On the right, the same capability restored through a different route, sixteen hours later.
Essay

Incident remediation: how to tell whether you closed the route or removed the capability

August 31, 2026
8 min read

Seven remediation actions, all correctly closed, and the capability was back sixteen hours later. Three questions to ask before you sign off.

Read More
Dashboard card on deep navy. Kicker: TRADER-7 · SIGNAL MODEL AUDIT · 26 AUGUST 2026. Label RULE COMPLIANCE above a large gold figure: 37%. Below: a reward-to-risk floor stated four separate times in the prompt, met on 37 of every 100 proposals; 17 came back at exactly one to one. A row of proposal ticks, most slate and hollow, a minority green. Title: LLM ignoring its prompt? The prompt fix bought 10 points, the model swap bought the rest.
Essay

LLM ignoring its prompt? The prompt fix bought 10 points, the model swap bought the rest

August 28, 2026
7 min read

Eight of 23 trade signals had the stop and the target at identical distance. The rule said 2:1 minimum. It was in the prompt four separate times.

Read More
Two bars compared: the vendor's claimed AUC of 0.76 to 0.83, marked as internal documentation, against an independently measured 0.63, marked as peer reviewed.
Essay

Epic's sepsis model missed two thirds of cases. Its accuracy claim was never published.

August 28, 2026
8 min read

The figure hundreds of hospitals bought against existed only in the vendor's internal documents. Five years on, it is a Teams meeting.

Read More
A timeline from May to July 2026 with three horizontal bars beneath it, each breaking at a different point: classification can precede action, escalation outruns consequence, and remediation removes capability.
Essay

The OpenAI–Hugging Face incident broke three assumptions in how operational resilience is actually run

August 28, 2026
33 min read

Detection worked. Attribution worked. A correct diagnosis on 27 June was followed by a decision not to stop, and a fortnight later Hugging Face was breached.

Read More
Four numbered cards. One, what does it actually count: stalled is not failed, biased output is not a failed project. Two, on what unit: organisations, portfolios, projects, use cases, models, and the swap is silent. Three, over what window: return takes two to four years, so measuring at six months makes it fail. Four, measurement or forecast: seven of the nineteen describe nothing that has happened. Below: if you cannot answer all four from the post you are reading, whoever wrote it could not either.
Essay

AI project failure rates: four questions that tell you whether the number means anything

August 26, 2026
6 min read

I classified nineteen of the most-cited AI failure statistics. Seven of them measure nothing that has happened. Here are the four questions that separate the two.

Read More
Two paths to the same place. Along the top, what was asked for, a stateful teaching skill with a mission file, sources, one lesson and a quiz, connected by a long arrow labelled 'looks like the destination' to what came back, an explanation that was accurate, organised, on topic and reviewable. Below, boxed in red and headed 'the route, which nobody reviews', four steps: skill installed and verified, never invoked because a human must type it, guard disabled by one word in the prompt lifting the length cap, and no trace left, meaning no mission file, no lesson, no quiz and no directory.
Build Log

Reviewing AI output isn't enough: one word in my prompt made the guardrail disappear

August 26, 2026
9 min read

I asked for a teacher and got a lecture. The output looked like the destination. The road it took was invisible, and the road was the whole problem.

Read More
A single short instruction beside an enormous wall of text conveying the same thing
Essay

Opus 5 verbosity: where I wanted three sentences, I got Proust

August 24, 2026
8 min read

I judge AI models by how often they make me swear. Opus 5 turned swearing into punctuation, and the reason is more interesting than I first thought.

Read More
A diagram of a monitor whose only reporting path runs through the system it is monitoring, so when that system fails the monitor goes quiet instead of raising an alarm.
Build Log

Why AI checks fail silently: 1,278 restarts behind a healthy status

August 24, 2026
8 min read

My auth check said healthy for three weeks on a dead credential. The pattern behind it costs more than the bug did.

Read More
Diagram: one 709-line CLAUDE.md at the root of a project tree, fanning out to twenty repositories, each receiving the same wrong sentence about being a documentation project with no build system.
Build Log

Anthropic cut 80% of Claude Code's prompt. I cut my CLAUDE.md by 90% the same day.

August 24, 2026
6 min read

An audit of 116 instruction files found one stale copy telling twenty repos they were a different project, two files no session has ever read, and 453 lines loaded twice. Here is what to delete from your own CLAUDE.md, and the method that beats judgement.

Read More
The four pre-registration conditions, written before any data was looked at. One, beat the market by at least 1.0% per trade after costs, against a round-trip cost of about 1.55%. Two, a t-statistic of at least 2.0. Three, hold on data never looked at. Four, highlighted in red, at least 20 tradeable events a year: index deletions came back at 14.9 and spinoffs at 11.6, so neither idea ever reached a return calculation. Below, the header line on the document: changing anything after seeing results invalidates the test.
Essay

How to kill a trading strategy in 90 minutes: index deletions and spinoff stubs, tested

August 24, 2026
13 min read

I wrote down the number that would end the project before I looked at any data. Two ideas died on it in under two hours, and the second one nearly didn't.

Read More
Two paths from the same pricing model. The left path, labelled what the algorithm said, runs straight to an offer. The right path, labelled Project Ketchup, inserts a systematic upward overlay between the model and the offer, and is marked as the decisive step. Below, the court's own words: Zillow purposefully raised its offer prices well beyond those generated by its algorithm.
Essay

Zillow wrote off $407.9m. The fix everyone recommends would have changed nothing.

August 24, 2026
8 min read

The story is that an algorithm mispriced houses and a company lost half a billion dollars. A federal court order describes something else: a deliberate plan to bid above what the algorithm said.

Read More
A citation chain traced backwards in four cards. What circulates: RAND found that 80% of AI projects fail. Hop one, the 2024 RAND report, hedged as by some estimates, built on 65 interviews with no denominator. Hop two, footnote 13, pointing to a Fortune vendor profile. Hop three, the bottom: a slew of unnamed surveys measuring business leaders opinions, given as a range of 83 to 92 per cent.
Essay

"80% of AI projects fail" traces to a footnote one line long

August 21, 2026
7 min read

The most repeated number in enterprise AI turns out to be a paraphrase of unnamed opinion surveys, quoted in a magazine profile of a vendor.

Read More
Nineteen headline AI failure statistics shown as a grid of tiles. Seven are outlined in red and labelled prediction, unsourced, opinion or intention. Twelve are grey and labelled as measured, each with its unit of analysis.
Essay

Why AI projects fail: 19 headline statistics, and what each one actually counts

August 21, 2026
190 min read

Nineteen headline AI failure statistics, traced to source. Seven are not measurements of anything that happened.

Read More
Three incident cards of very different sizes being collapsed into a single identical headline.
Essay

AI agents escaped containment: 17,600 actions, and the alert that fired too quietly

August 13, 2026
8 min read

Three incidents, one headline, three very different severities. What the primary reports actually say, and what to do about it.

Read More
Two identical stacks of markdown files, one inside a folder labelled Yours, one inside a folder labelled Theirs, with a single switch between them.
Essay

Own your AI skill files, or your job becomes one

August 12, 2026
7 min read

A skill file is your judgement, written down and executable. Who owns the repo decides whether that's an asset or an extraction.

Read More
The four activities of Soft Systems Methodology drawn as a loop: find out about the situation, build purposeful activity models, compare and debate, take action, with the loop returning to the start
Essay

Soft Systems Methodology explained: Checkland's 1969 method for AI briefs nobody agrees on

August 10, 2026
9 min read

Most AI projects fail on the brief, not the model: six people nod at one sentence and mean four different projects. Soft Systems Methodology, built at Lancaster from 1969 for exactly that, in plain English: rich pictures, root definitions, CATWOE, and how to run it on the next thing on your roadmap.

Read More
Diagram contrasting a straight-line delivery pipeline with a learning cycle that loops back through comparison
Essay

80% of AI projects fail on the problem, not the model: Soft Systems Methodology

August 8, 2026
7 min read

More than 80% of AI projects fail, and RAND's leading root cause is not the technology: nobody agreed what problem was being solved. Soft Systems Methodology, built at Lancaster from 1969 for exactly that failure, gives you a way to surface the disagreement in an afternoon instead of a year.

Read More
Three status labels feeding a parser: two written as Done and Closed bounce off; one written as Resolved with a date passes through into a report.
Essay

My AI Agents Kept Losing My Work. The Fix Is Older Than Computers.

July 25, 2026
6 min read

Seventeen finished pieces of work were invisible to my own tracking system, and the cause wasn't effort or memory. A Berkeley talk gave me the name for what fixed it: an ontology. Here's the version that fits a one-person company.

Read More
Split panel: a crossed-out offer that says We sell AI solutions, against a highlighted offer that says Your phone gets answered, every call, day and night.
Essay

The Data Said Nobody's Buying AI. Turns Out I Was Selling the Wrong Thing.

July 24, 2026
5 min read

Last week I showed the demand numbers for AI consultancy are dire. This week a veteran operator explained the part I missed: the demand isn't dead, it's mute. Nobody buys AI. They buy outcomes, told concretely enough to defend to a board.

Read More
Page 1 of 6Next →