Skip to main content

Become more valuable, not less, as AI accelerates.

Business continuity and incident management in banking since 2000, and building with AI in the open now. I test what's claimed about AI against the primary sources and publish what I find: long papers, each backed by a series of shorter articles you can use.

Jamie WattersOperational resilience and AI delivery practitioner4 published books17 products built with AI

Get it in your inbox

One weekly email. Real numbers, the failures before the wins.

Research programmes

A long paper on each subject, with a series of shorter articles that came off it. Every figure gets chased back to where it came from, and where the trail stops early the piece says so.

Paper

Test an AI model before you rely on it: six mistakes I made testing Jev

23 September 2026 · 22 min read

One run said Jev beat GPT-5.4 by 11.1 points. Five runs could not tell them apart. How to test an AI model before you rely on it, from the six mistakes behind that headline.

The series, in reading order

  1. Jev is not an LLM. One run had it beating GPT-5.4; five reruns couldn't tell them apart.20 September 2026
  2. GPT-5.4 beats Jev on a confidence threshold, once tested at its own values.21 September 2026
  3. Paying your AI to think longer didn't fix its mistakes. It cost up to 4.2x more.22 September 2026
  4. I said Jev beat GPT-5.4. Five reruns can't tell them apart, at a fiftieth of the cost.23 September 2026

Who this is for

For anyone embracing AI to build a better life and stay relevant as the world changes. That's the builder or founder trying to make sense of it without drowning in hype, the career professional wondering what survives, the late starter, and the solo operator who's never written code.

I'm not an AI guru selling a course of recycled ideas, and I'm not a consultant who's never shipped anything. I'm the man in the water: building things with these tools, finding out what actually holds up, and telling you the truth about it before you spend your own time finding out.

There's a quieter thread here too. As AI gets better at thinking for you, the skill that matters most is thinking for yourself: keeping your own judgement instead of handing it over. I'm working that out in public alongside everything else.

How the writing works

Most of what I write is ad hoc: something I learned building with AI, a piece of philosophy, a writing lesson, something interesting I came across. When a subject deserves more, I go deeper, and that becomes a long paper supported by a series of shorter, usable articles.

One habit runs through all of it: I check what's being reported against the primary sources. The reported version is shaped by commercial interest or an existing narrative. The gap between that and what the sources actually show is usually the most interesting thing in the piece.

The tools are part of the same habit. I build things to learn, not to sell, give the code away, and write up what worked, what didn't, and why. When something stops earning its place, I kill it in public and tell you what it cost me.

Why it matters

Hundreds of millions of people now get their understanding of the world from a handful of similarly-trained AI models. The answers come back confident, reasonable, and identical for everyone. When the machine hands the whole crowd the same play, the play stops paying.

So the scarce thing, the only defensible thing, is the ability to keep your own mind: real judgement, grounded in real depth, from someone actually doing the work. That's what I'm building here. Not another feed of AI takes. A field report you can verify.

Why you might believe me

  • Business continuity, crisis and incident management in banking since 2000, with the US regulatory reviews of those programmes (OCC, NFA and the Federal Reserve) closed with no issues raised
  • Started in data centres, then as a systems programmer writing OS code in assembler, before working across IT and into IT project, program and senior management
  • 4 published books, including The Business Continuity Management Desk Reference, the best-selling book on business continuity in the world for its first few years
  • 17 products built with AI, across search, trading, monitoring, and estate tooling, several since retired
  • Creator of Efformism, a philosophical framework, and of innovations on The Headless Way, a contemplative practice
  • Code open for anyone to inspect; products killed in public when they don't work

The field report

The latest posts, written as I go. What's working, what isn't, and what it cost me.

A card reading: that chart where one AI beats another? Run it again and the winner can vanish. Mine did. Ask before you pick the tool: was it run more than once, who marked the answers, did the two really disagree?
Essay

Test an AI model before you rely on it: six mistakes I made testing Jev

23 September 2026
22 min read

One run said Jev, TypeSafe's new model, beat GPT-5.4 by 11.1 points. Five runs couldn't tell them apart. The six mistakes behind that headline, and a checklist to run before you rely on any AI comparison, mine included.

Read More
A results card reading: I said Jev beat GPT-5.4, five reruns can't tell them apart. Jev's lead over GPT-5.4 on the hard claims was 11.1 points in one run and 3.3 points averaged over five runs, which is 0.6 of one claim and not significant.
Build Log

I said Jev beat GPT-5.4. Five reruns can't tell them apart, at a fiftieth of the cost.

23 September 2026
13 min read

I published that Jev, TypeSafe's new model, beat GPT-5.4 by 11.1 points on hard fact-checks. Rerun five times, the two can't be told apart, Jev costs a fiftieth as much, and it lets more wrong claims through. Plus the five-run check.

Read More
A results card reading: paying your AI to think longer didn't fix its mistakes. Up to 4.2x extra cost per answer for GPT-5.4, Sonnet 5 and Gemini 3.1 Pro, and no net gain on the hard claims for any of the three, one run per model.
Build Log

Paying your AI to think longer didn't fix its mistakes. It cost up to 4.2x more.

22 September 2026
14 min read

I gave three AI models room to think before answering, the fix everyone assumes will work. None got more reliable. It cost 2.4 to 4.2 times more per answer to find that out, and Jev, the cheap specialist model, still caught more of the mistakes.

Read More

Tools I built to learn

Built with AI to find out what holds up, open to inspect, and killed in public when they stop earning their place.

Executor File

Everything your executor will need, in one encrypted file.

Beta

Everything your executor will need, in one encrypted file. A self-hosted register of accounts, assets and liabilities, with no credentials stored and no service to die.

Bash/POSIX shawkage (encryption)ssss (Shamir)+2

AI Search Mastery

The static hub for the AI search work: the MASTERY-AI framework, guides and articles.

Live

The static hub for the AI search work: the MASTERY-AI framework, guides and articles. The paid tools it once fronted are gone.

Static HTMLCSSJavaScript

The deal

If I recommend it, I'm using it.

If it failed, you'll hear about the failure before the win.

My code is open for you to check.

Every product I kill gets a public post-mortem.

That's the whole deal. No funnel, no secret sauce, no "DM me to scale."

Questions people actually ask

Do you build all of this yourself?
Yes: me plus AI. Most of the code is written with Anthropic's Claude and reviewed, tested and shipped by me. That working method is the subject of the site, not a secret behind it.
Is the code really open?
The product code is public on GitHub under TheWayWithin, open for anyone to inspect. When a product stops earning its place I kill it in public and write the post-mortem.
What do subscribers actually get?
A weekly digest of the field reports: what I built, what worked, what failed and what it cost. You confirm by email, no spam, unsubscribe any time.
Are you selling a course or consulting?
Not through this site. The writing is free, and the newsletter is the writing by email. The tools stand on their own and say on their own pages what they charge for. The work I get paid for is delivery, not advice from the sidelines: operational resilience and AI deployment inside regulated firms. If that is what you are after, email me.

I'm in the water, riding this wave in real time, in code and in writing.

If you'd rather learn from someone actually doing the work than someone selling theory from the beach, stay a while. I'll tell you what's working before you waste the time finding out yourself.

Get the next one in your inbox

I build with AI in the open and write up what held and what didn't. Real numbers, the failures before the wins.