Skip to main content

The Journey

Field reports from building with AI in public. What's working, what isn't, and what it cost me. Open numbers, the failures before the wins.

Subscribe via RSS
Dashboard card on deep navy. Kicker: TRADER-7 · LIVE MODEL SWAP · 22 July 2026. Large gold figure: 4.7 arrow 4.8. Below: both decision models swapped live with one env var; 78 real prompts replayed first, zero regression, paper account still plus $1,234. A row of even green dots crossed by a thin slate SWAP marker, the dots unbroken on both sides. Title: I swapped the model and nothing happened.
Build Log

I swapped the model and nothing happened

August 9, 2026
5 min read

Swapping the two LLMs behind an autonomous trading bot from Opus 4.7 to 4.8, live, without it missing a beat. The trick was replaying 78 real prompts first.

Read More
AI projects fail on the problem, not the model. Below: RAND's leading root cause is that nobody agreed what problem was being solved, one brief yields four root definitions, and the fourth is never written down.
Essay

80% of AI projects fail on the problem, not the model: Soft Systems Methodology

August 8, 2026
7 min read

More than 80% of AI projects fail, and RAND's leading root cause is not the technology: nobody agreed what problem was being solved. Soft Systems Methodology, built at Lancaster from 1969 for exactly that failure, gives you a way to surface the disagreement in an afternoon instead of a year.

Read More
Where I wanted three sentences, I got Proust. Below: a crash stops you and costs minutes, plausible wrong output stops nothing and costs the afternoon, and only one of them is being counted.
Essay

Opus 5 verbosity: where I wanted three sentences, I got Proust

August 6, 2026
8 min read

I judge AI models by how often they make me swear. Opus 5 turned swearing into punctuation, and the reason is more interesting than I first thought.

Read More
My auth check said healthy for three weeks on a dead credential. Below: the check ran and passed every time, the oracle was answering the wrong question, and the pattern costs more than the bug did.
Build Log

Why AI checks fail silently: 1,278 restarts behind a healthy status

August 3, 2026
8 min read

My auth check said healthy for three weeks on a dead credential. The pattern behind it costs more than the bug did.

Read More
Anthropic cut 80% of the prompt and I cut my instructions by 90%. Below: one stale copy telling twenty repos they were a different project, two files no session has ever read, and 453 lines loaded twice.
Build Log

Anthropic cut 80% of Claude Code's prompt. I cut my CLAUDE.md by 90% the same day.

August 3, 2026
6 min read

An audit of 116 instruction files found one stale copy telling twenty repos they were a different project, two files no session has ever read, and 453 lines loaded twice. Here is what to delete from your own CLAUDE.md, and the method that beats judgement.

Read More
Kastrup's cosmic mind has no plan. Below: it is instinctive rather than deliberating, nothing is being worked towards, and it is not looking back at you.
Essay

Bernardo Kastrup's cosmic mind has no plan: what analytic idealism actually claims

August 1, 2026
10 min read

The most rigorous idealist alive gives the non-duality audience far less than they want: a universal consciousness that is instinctive, has no plan, and is not looking back at you. What he actually claims, why he says it cannot be proven, and the one part that survives whoever turns out to be right.

Read More
Three status labels feeding a parser: two written as Done and Closed bounce off; one written as Resolved with a date passes through into a report.
Essay

My AI Agents Kept Losing My Work. The Fix Is Older Than Computers.

July 25, 2026
6 min read

Seventeen finished pieces of work were invisible to my own tracking system, and the cause wasn't effort or memory. A Berkeley talk gave me the name for what fixed it: an ontology. Here's the version that fits a one-person company.

Read More
Split panel: a crossed-out offer that says We sell AI solutions, against a highlighted offer that says Your phone gets answered, every call, day and night.
Essay

The Data Said Nobody's Buying AI. Turns Out I Was Selling the Wrong Thing.

July 24, 2026
5 min read

Last week I showed the demand numbers for AI consultancy are dire. This week a veteran operator explained the part I missed: the demand isn't dead, it's mute. Nobody buys AI. They buy outcomes, told concretely enough to defend to a board.

Read More
Conference-thumbnail style graphic: a bordered jamiewatters.work badge, an orange markdown-logo square with the word Markdown, and giant white text reading Go find the customers
Essay

Garry Tan Just Described My Setup. Then I Noticed What He Left Out.

July 21, 2026
7 min read

The YC president gave a talk describing, almost item for item, the AI system I already run. The validation lasted about ten minutes. Then I noticed the word neither of us was saying.

Read More
I got phished from my own verified domain. Below: the key leaked by me in public, the domain verified and now the weapon, and it happened faster than any review cycle.
Essay

I Leaked an API Key and It Turned My Own Domain Into a Phishing Weapon

July 21, 2026
6 min read

I got phished from my own verified domain, using an API key I'd leaked myself: what actually leaked, how fast it happened, and what I changed.

Read More
The map is a file. It outlasts every company that might have owned it. Below: a friend said it should be a business, I assessed it honestly and decided it was not, then built it anyway.
Essay

A friend said it was a business. I decided it was not, and built it anyway.

July 19, 2026
3 min read

A friend said it should be a business. I assessed it honestly, decided it was not, and built it anyway because it still needed to exist.

Read More
A simple data panel: only 8 to 18 percent of small firms pay for any AI, median spend about $28 a month, and only about 8 percent hire a consultant for AI.
Essay

Everyone's selling AI consultancy. The data says nobody's buying.

July 19, 2026
4 min read

A $1,000-an-hour AI consultancy playbook is doing the rounds. The demand data, and a friend who sells everything except this, disagree.

Read More
Two workbenches building the same object, one cluttered with scaffolding, one spare, both producing an identical result.
Build Log

I built the same tool two ways to settle an argument

July 16, 2026
5 min read

Spec-first or goal-first? I didn't believe either side, so I built the same tool both ways with one spec and one model. Both worked. The bill didn't tie.

Read More
Dashboard card on Deep Space navy background. Kicker: UNATTENDED, 22 JUN to 3 JUL, 11 DAYS. Large gold number: plus $132.00. Below: 5 losses, 1 winner (BTC plus $49.72), loss-streak pause fired at 5 and 7, no emergency halts. Sparkline: five red dots stepping down, one green dot jumping up. Title: Eleven days I didn't touch it.
Build Log

Eleven days I didn't touch it

July 4, 2026
5 min read

Eleven days of unattended trading. Five losing trades, one winner covered them, +$132 net. What the system did right, what drifted, what I fixed on return.

Read More
On the left, a lone clamp gripping a single tool; on the right, an organised wall of specialist tools around a small instrument sealed under glass with a wax seal.
Build Log

How to trust AI-written code you didn't read line by line

June 22, 2026
5 min read

AI can write working code in minutes. The hard part is trusting it without reading every line. Here's what actually makes AI-built code trustworthy.

Read More
A craftsman's workbench with hand tools and wood shavings beside a brass measuring gauge sealed under a glass dome with a wax seal
Build Log

Everyone's selling autonomous AI agents. I built one, and the real lesson was 20 years old.

June 21, 2026
4 min read

An AI agent spent an afternoon optimising one of my apps. It tried a change, measured it, kept the one that was genuinely faster, reverted three that weren...

Read More
The privacy trap is real, just not the one everyone warned me about. Below: six AIs said it uploads your code, it does not, and the real gotcha is smaller and easy to miss.
Build Log

I Ran Graphify. The Privacy Trap Is Real, Just Not the One Everyone Warned Me About

June 20, 2026
7 min read

Six AIs told me Graphify uploads your code. I checked by running it, and the actual gotcha is smaller, quieter, and easy to miss.

Read More
Claude keeps killing the good bit by turning it into a multiple choice. Below: an open question becomes three options, the best answer was usually the fourth, and once you see why you can stop it.
Build Log

The Three Options Trap: Why Claude Keeps Killing the Good Bit

June 18, 2026
7 min read

Claude has a habit of collapsing an open conversation into three multiple-choice options, and once you see why it does it, you can stop it doing it to you.

Read More
Six AIs handed me the same lie, in unison. Below: unanimous reads as corroboration, but it was one claim repeated, so I check the source before I publish.
Build Log

I Almost Told You a Lie. Six AIs Handed It to Me, in Unison.

June 17, 2026
7 min read

A field note from researching Graphify: the real privacy model, the honest token numbers, and why I check the source before I publish.

Read More
Three-period performance comparison: Pre-Sprint 120 at 41.2% win rate +$964, Sprint 120-136 at 24.4% win rate -$217, Post-Sprint 136 at 47.6% win rate +$321.
Build Log

Same system, three different traders

May 26, 2026
4 min read

Same code, same market, three time periods. One made $964, one lost $217, one's up $321. The difference wasn't the algorithm.

Read More