The Journey
Most posts here are ad hoc: what I learned building with AI that week, some philosophy, a writing lesson, something I came across.
When a subject deserves more than a post it becomes a long paper, with a series of shorter articles behind it.
The papers
- The OpenAI–Hugging Face incident broke three assumptions in how operational resilience is actually run
- Why AI projects fail: 19 headline statistics, and what each one actually counts
- Test an AI model before you rely on it: six mistakes I made testing Jev
All posts, latest first

If AI saves you time, why aren't you being paid more for it?
In a study of 25,000 Danish workers, people saving time with AI say they expect to spend it on other work, and their pay hasn't pulled ahead of non-users' by more than 2% on average. Who decides what saved time becomes, and a log to track your own.

An AI company wants to buy your firm? Ask it to prove the AI works first
AI-backed companies worth billions are buying accountants and IT firms. If one calls you, here is what proof its AI works would look like, who has shown any of it, and what to ask before you sell.

Amazon's 30,000 job cuts weren't proof AI is taking jobs. What its boss says is coming is closer to it
Amazon said AI was why it had to move faster, not who did the work. Its boss expects AI to shrink its corporate workforce later. Here's the record, in date order.

Automating with AI? Australia cut one human check from its welfare debt system and unlawfully billed 433,000 people
Robodebt cut a slow human check to go faster. Officers then redid 919 of its debts by hand and marked 29 per cent wrong, and the volume went up anyway. Five things to do before AI takes over a process.

OpenAI's Texas data centre is built to barely use water. The power plants it runs on use 300 times more water than it does
The company building OpenAI's Stargate campus in Abilene says a 200 megawatt block uses 910,000 gallons of water a year on its website, and about 2 million in…

AI water use: the bottle of water per 20 ChatGPT questions was withdrawn by its own authors in 2023. It's still quoted in 2026
The most quoted AI water number, a bottle per 20 ChatGPT questions, was withdrawn by its own authors in October 2023 and is still being repeated. Every AI…

Jev is a narrow tool with four clear uses. Greg Isenberg's $10m framing is not one of them.
Greg Isenberg's episode on Jev promised $10m businesses. His own guest says use it for routing and keep it away from your money. Here is each claim against the data, and the four conditions under which the tool earns its place.

Test an AI model before you rely on it: six mistakes I made testing Jev
One run said Jev, TypeSafe's new model, beat GPT-5.4 by 11.1 points. Five runs couldn't tell them apart. The six mistakes behind that headline, and a checklist to run before you rely on any AI comparison, mine included.

Jev, a cheap AI, did two thirds of my filing: same result as Claude, 63% cheaper
Jev, a cheap AI, handled two thirds of my filing and passed the rest to Claude. Same result as Claude alone, 63% cheaper. The catch: the hand-off isn't a check, because Claude makes the same mistakes.

I said Jev beat GPT-5.4. Five reruns can't tell them apart, at a fiftieth of the cost.
I published that Jev, TypeSafe's new model, beat GPT-5.4 by 11.1 points on hard fact-checks. Rerun five times, the two can't be told apart, Jev costs a fiftieth as much, and it lets more wrong claims through. Plus the five-run check.

Paying your AI to think longer didn't fix its mistakes. It cost up to 4.2x more.
I gave three AI models room to think before answering, the fix everyone assumes will work. None got more reliable. It cost 2.4 to 4.2 times more per answer to find that out, and Jev, the cheap specialist model, still caught more of the mistakes.

GPT-5.4 beats Jev on a confidence threshold, once tested at its own values.
I published a wrong finding last week: that GPT-5.4's confidence score has no usable threshold. It does, and swept over its own values it beats Jev at every…

Jev is not an LLM. On the hard half it beat GPT-5.4, Gemini 3.1 Pro and Sonnet 5, at a fiftieth of the cost.
Jev does not generate text. You give it a question with defined answers and it hands back one of them plus a probability. On 42 citation checks it tied the…

The file your AI agent wrote for itself is making it worse. Ten minutes fixes it.
Somewhere in your project is a file your agent wrote, that your agent now reads as fact. Nothing in between checks it, so the good and the bad accumulate…

AI self-improvement fails at the scorer, in every case I could check
Google shipped a working self-improvement technique four days ago. One sentence in the paper explains why it works: the evaluator is frozen and sits outside the loop. Then look at the failures. Pseudo-labels that decay, a reward that gets gamed, a detector that gets bypassed, skills nobody scored. Different systems, same broken part. And if you run AI agents, you are already running the ungoverned version of this.

AI arms race game theory: the game we're in turns on an unmeasured belief
Everyone calls the AI race a prisoner's dilemma. The one formal model built to test that gets two worlds, and which one you're in depends on both sides…

Anthropic's alignment lead put extinction above 10%: ask where the plan is, not whether the number is right
The argument went to the percentage. The same three sentences say there is no plan for the risk being priced. Here is what has to sit next to a likelihood…

AI agent security: find the file your agents can write and every agent reads
In one research harness, a payload sitting in the file that loads into the next agent's system prompt infected that agent 55% of the time. The same payload in…

Checking the source is not enough: I traced a number to a scientific paper and found a Teams meeting
Everyone tells you to check the source. Here is a case where a careful person does exactly that, reads the peer-reviewed paper, and still comes away believing…

The credibility test we run backwards: ask who they were paying
One researcher quit and forfeited unvested equity to warn about AI. Another said the same thing for free and kept his job. The test that sorts them is not what it cost but who they were paying, and this piece corrects the version of it I published a day earlier.