Skip to main content

The Journey

Most posts here are ad hoc: what I learned building with AI that week, some philosophy, a writing lesson, something I came across.

Subscribe via RSS

All posts, latest first

Social card. Headline: Workers expect to put the time AI saves into other work. Their pay hasn't pulled ahead of non-users' by more than 2% on average. Below: who decides what saved time becomes. Source: Humlum and Vestergaard, March 2026, a study of 25,000 workers in Denmark.
Essay

If AI saves you time, why aren't you being paid more for it?

7 October 2026
12 min read

In a study of 25,000 Danish workers, people saving time with AI say they expect to spend it on other work, and their pay hasn't pulled ahead of non-users' by more than 2% on average. Who decides what saved time becomes, and a log to track your own.

Read More
Social card. Headline: Companies worth billions are buying accountants to run them with AI. Below: what would prove it works, and who has shown it so far.
Essay

An AI company wants to buy your firm? Ask it to prove the AI works first

6 October 2026
10 min read

AI-backed companies worth billions are buying accountants and IT firms. If one calls you, here is what proof its AI works would look like, who has shown any of it, and what to ask before you sell.

Read More
Social card. Headline: Amazon hasn't said AI did the jobs it cut. Below: its chief executive expects AI to reduce the corporate workforce in the next few years.
Essay

Amazon's 30,000 job cuts weren't proof AI is taking jobs. What its boss says is coming is closer to it

5 October 2026
11 min read

Amazon said AI was why it had to move faster, not who did the work. Its boss expects AI to shrink its corporate workforce later. Here's the record, in date order.

Read More
Social card: They checked the computer's debts by hand. 29% were wrong. The volume went up anyway. Cut for speed: the officer who checked whether a mismatch was real. The pace: 20,000 a year became 20,000 a week, by the minister's count. The check: officers redid 919 debts by hand and marked 29% wrong. Robodebt, Australia, 2016, rules-based, not AI.
Essay

Automating with AI? Australia cut one human check from its welfare debt system and unlawfully billed 433,000 people

4 October 2026
9 min read

Robodebt cut a slow human check to go faster. Officers then redid 919 of its debts by hand and marked 29 per cent wrong, and the volume went up anyway. Five things to do before AI takes over a process.

Read More
Card headed 'Built to barely use water: one 200 megawatt block of the Stargate campus in Abilene, Texas.' Three rows. The builder's website: 910,000 gallons a year. The builder's own report: about 2 million gallons a year, about 20,700 litres a day, three Kansas restaurants. The power plants making its electricity, on the Texas grid average: about 6.2 million litres a day, and a water-cooled site pays this too. Footer: Crusoe 2025 impact report and web page, VanSchenkhof 2011, Li and colleagues 2025 using a World Resources Institute 2020 factor.
Essay

OpenAI's Texas data centre is built to barely use water. The power plants it runs on use 300 times more water than it does

2 October 2026
10 min read

The company building OpenAI's Stargate campus in Abilene says a 200 megawatt block uses 910,000 gallons of water a year on its website, and about 2 million in…

Read More
Card headed 'The most quoted AI water number was withdrawn by its own authors in October 2023', with two dated quotes from the same paper: 6 April 2023, 'ChatGPT needs to drink a 500ml bottle of water for a simple conversation of roughly 20-50 questions and answers', and 25 October 2023, 'GPT-3 needs to drink (i.e., consume) a 500ml bottle of water for roughly 10-50 responses'. Footer: the first sentence is the one still quoted in 2026; source arXiv 2304.03271, versions 1 and 2.
Essay

AI water use: the bottle of water per 20 ChatGPT questions was withdrawn by its own authors in 2023. It's still quoted in 2026

29 September 2026
9 min read

The most quoted AI water number, a bottle per 20 ChatGPT questions, was withdrawn by its own authors in October 2023 and is still being repeated. Every AI…

Read More
A card quoting Greg Isenberg: $10 million businesses, $100 million businesses. Beside it, his own guest saying to use it for routing only. The expensive queue the episode claims to unlock costs $1.29 a month on the old model for a thousand enquiries, so nothing was rationed. A narrow tool with four clear uses.
Essay

Jev is a narrow tool with four clear uses. Greg Isenberg's $10m framing is not one of them.

24 September 2026
10 min read

Greg Isenberg's episode on Jev promised $10m businesses. His own guest says use it for routing and keep it away from your money. Here is each claim against the data, and the four conditions under which the tool earns its place.

Read More
A card reading: that chart where one AI beats another? Run it again and the winner can vanish. Mine did. Ask before you pick the tool: was it run more than once, who marked the answers, did the two really disagree?
Essay

Test an AI model before you rely on it: six mistakes I made testing Jev

23 September 2026
22 min read

One run said Jev, TypeSafe's new model, beat GPT-5.4 by 11.1 points. Five runs couldn't tell them apart. The six mistakes behind that headline, and a checklist to run before you rely on any AI comparison, mine included.

Read More
A card reading: A cheap AI did two thirds of my filing. Same result as Claude, 63% cheaper. Jev passing unsure notes to Claude 86.7% at $0.0034 a note; Claude alone 86.7% at $0.0091 a note; keyword search with no AI 76.0%. But it kept Claude's mistakes.
Build Log

Jev, a cheap AI, did two thirds of my filing: same result as Claude, 63% cheaper

23 September 2026
12 min read

Jev, a cheap AI, handled two thirds of my filing and passed the rest to Claude. Same result as Claude alone, 63% cheaper. The catch: the hand-off isn't a check, because Claude makes the same mistakes.

Read More
A results card reading: I said Jev beat GPT-5.4, five reruns can't tell them apart. Jev's lead over GPT-5.4 on the hard claims was 11.1 points in one run and 3.3 points averaged over five runs, which is 0.6 of one claim and not significant.
Build Log

I said Jev beat GPT-5.4. Five reruns can't tell them apart, at a fiftieth of the cost.

23 September 2026
13 min read

I published that Jev, TypeSafe's new model, beat GPT-5.4 by 11.1 points on hard fact-checks. Rerun five times, the two can't be told apart, Jev costs a fiftieth as much, and it lets more wrong claims through. Plus the five-run check.

Read More
A results card reading: paying your AI to think longer didn't fix its mistakes. Up to 4.2x extra cost per answer for GPT-5.4, Sonnet 5 and Gemini 3.1 Pro, and no net gain on the hard claims for any of the three, one run per model.
Build Log

Paying your AI to think longer didn't fix its mistakes. It cost up to 4.2x more.

23 September 2026
14 min read

I gave three AI models room to think before answering, the fix everyone assumes will work. None got more reliable. It cost 2.4 to 4.2 times more per answer to find that out, and Jev, the cheap specialist model, still caught more of the mistakes.

Read More
A card reading: what TypeSafe's own 0.8 auto-accept threshold buys you, and that a model's confidence threshold has to be tested at its own numbers, not a grid chosen in advance
Build Log

GPT-5.4 beats Jev on a confidence threshold, once tested at its own values.

21 September 2026
26 min read

I published a wrong finding last week: that GPT-5.4's confidence score has no usable threshold. It does, and swept over its own values it beats Jev at every…

Read More
A results table comparing Jev with GPT-5.4, Claude Sonnet 5 and Gemini 3.1 Pro across 42 citation checks, split into an easy set and a hard set
Build Log

Jev is not an LLM. On the hard half it beat GPT-5.4, Gemini 3.1 Pro and Sonnet 5, at a fiftieth of the cost.

20 September 2026
18 min read

Jev does not generate text. You give it a question with defined answers and it hands back one of them plus a probability. On 42 citation checks it tied the…

Read More
Social card reading: Your agent wrote a file. Your agent now reads that file as instruction. Nothing in between checked it.
Essay

The file your AI agent wrote for itself is making it worse. Ten minutes fixes it.

19 September 2026
7 min read

Somewhere in your project is a file your agent wrote, that your agent now reads as fact. Nothing in between checks it, so the good and the bad accumulate…

Read More
Social card reading: Seven papers on AI that improves itself. Two ways it goes wrong, and both are the same part. The thing doing the scoring breaks, or nothing is scoring what you just added. The one technique that works freezes its evaluator outside the loop.
Essay

AI self-improvement fails at the scorer, in every case I could check

18 September 2026
21 min read

Google shipped a working self-improvement technique four days ago. One sentence in the paper explains why it works: the evaluator is frozen and sits outside the loop. Then look at the failures. Pseudo-labels that decay, a reward that gets gamed, a detector that gets bypassed, skills nobody scored. Different systems, same broken part. And if you run AI agents, you are already running the ungoverned version of this.

Read More
Social card reading: a prisoner's dilemma in one world, a coordination problem in the other, and what decides which is a shared belief the model assumes rather than measures.
Essay

AI arms race game theory: the game we're in turns on an unmeasured belief

16 September 2026
14 min read

Everyone calls the AI race a prisoner's dilemma. The one formal model built to test that gets two worlds, and which one you're in depends on both sides…

Read More
Card reading: one number, no plan. A likelihood with no method, no owner and no controls, and the same three sentences saying there is no plan.
Essay

Anthropic's alignment lead put extinction above 10%: ask where the plan is, not whether the number is right

15 September 2026
11 min read

The argument went to the percentage. The same three sentences say there is no plan for the risk being priced. Here is what has to sit next to a likelihood…

Read More
Social card reading: seven hold no write path, two hold a shell under a read-only label, and all eleven can hand the write to an agent that has one. The door I left open is the file I write my own lessons into.
Build Log

AI agent security: find the file your agents can write and every agent reads

15 September 2026
16 min read

In one research harness, a payload sitting in the file that loads into the next agent's system prompt infected that agent 55% of the time. The same payload in…

Read More
Social card reading: Check the source, they said. I did. The paper cited a Teams meeting. A number reaches you at full strength. Whatever qualified it does not.
Essay

Checking the source is not enough: I traced a number to a scientific paper and found a Teams meeting

14 September 2026
9 min read

Everyone tells you to check the source. Here is a case where a careful person does exactly that, reads the peer-reviewed paper, and still comes away believing…

Read More
Card reading: one paid to say it, the other paid nothing, and the cost is only half the sum.
Essay

The credibility test we run backwards: ask who they were paying

14 September 2026
10 min read

One researcher quit and forfeited unvested equity to warn about AI. Another said the same thing for free and kept his job. The test that sorts them is not what it cost but who they were paying, and this piece corrects the version of it I published a day earlier.

Read More
Page 1 of 13Next →