AI
Working with AI, building with it, and thinking about where it goes.

80% of AI projects fail on the problem, not the model: Soft Systems Methodology
More than 80% of AI projects fail, and RAND's leading root cause is not the technology: nobody agreed what problem was being solved. Soft Systems Methodology, built at Lancaster from 1969 for exactly that failure, gives you a way to surface the disagreement in an afternoon instead of a year.

Opus 5 verbosity: where I wanted three sentences, I got Proust
I judge AI models by how often they make me swear. Opus 5 turned swearing into punctuation, and the reason is more interesting than I first thought.

Why AI checks fail silently: 1,278 restarts behind a healthy status
My auth check said healthy for three weeks on a dead credential. The pattern behind it costs more than the bug did.

Anthropic cut 80% of Claude Code's prompt. I cut my CLAUDE.md by 90% the same day.
An audit of 116 instruction files found one stale copy telling twenty repos they were a different project, two files no session has ever read, and 453 lines loaded twice. Here is what to delete from your own CLAUDE.md, and the method that beats judgement.

My AI Agents Kept Losing My Work. The Fix Is Older Than Computers.
Seventeen finished pieces of work were invisible to my own tracking system, and the cause wasn't effort or memory. A Berkeley talk gave me the name for what fixed it: an ontology. Here's the version that fits a one-person company.

The Data Said Nobody's Buying AI. Turns Out I Was Selling the Wrong Thing.
Last week I showed the demand numbers for AI consultancy are dire. This week a veteran operator explained the part I missed: the demand isn't dead, it's mute. Nobody buys AI. They buy outcomes, told concretely enough to defend to a board.

Garry Tan Just Described My Setup. Then I Noticed What He Left Out.
The YC president gave a talk describing, almost item for item, the AI system I already run. The validation lasted about ten minutes. Then I noticed the word neither of us was saying.

I Leaked an API Key and It Turned My Own Domain Into a Phishing Weapon
I got phished from my own verified domain, using an API key I'd leaked myself: what actually leaked, how fast it happened, and what I changed.

Everyone's selling AI consultancy. The data says nobody's buying.
A $1,000-an-hour AI consultancy playbook is doing the rounds. The demand data, and a friend who sells everything except this, disagree.

A goal prompt beat my build framework on a real job
Same spec, same model, two methods: a seven-part goal prompt against a full agent framework. The lean prompt tied on correctness and won on time, cost, and clutter. Here are the numbers.

How to trust AI-written code you didn't read line by line
AI can write working code in minutes. The hard part is trusting it without reading every line. Here's what actually makes AI-built code trustworthy.

Everyone's selling autonomous AI agents. I built one, and the real lesson was 20 years old.
An AI agent spent an afternoon optimising one of my apps. It tried a change, measured it, kept the one that was genuinely faster, reverted three that weren...

I Ran Graphify. The Privacy Trap Is Real, Just Not the One Everyone Warned Me About
Six AIs told me Graphify uploads your code. I checked by running it, and the actual gotcha is smaller, quieter, and easy to miss.

The Three Options Trap: Why Claude Keeps Killing the Good Bit
Claude has a habit of collapsing an open conversation into three multiple-choice options, and once you see why it does it, you can stop it doing it to you.

I Almost Told You a Lie. Six AIs Handed It to Me, in Unison.
A field note from researching Graphify: the real privacy model, the honest token numbers, and why I check the source before I publish.

Building David, Not a Smaller Goliath
How a trading agent reveals the game retail can actually win — by playing where the funds can't or won't go, and learning to love the play.

BOS-AI had twelve XML files holding nothing
I checked the XML files BOS-AI used to store 'institutional memory' for thirty agents. Every tag was PLACEHOLDER_X. So I deleted all twelve.

Half of agent-11 was scaffolding the platform now provides
I shipped v6 of agent-11 last month. The biggest single change was deletion. The deployed CLAUDE.md went from roughly 250 lines to under 80. The MCP prof...

Now the Real Data Starts
Seven months of building bought me a system that finally tells the truth about itself. The data those seven months produced is mostly unreadable. The clock just started, and not a moment too soon.
