Skip to main content

An AI can invent a belief you never held from your own notes. Mine did, for 18 days.

Jamie Watters

Operational resilience and AI delivery practitioner. Technology since 1985.

Published: 14 September 20267 min read
#belief-audit#epistemics#truth-rule#thinking
Title card reading: an AI can invent a belief you never held from your own notes, mine did, for 18 days. Below it, two boxes: recorded as, a claim the agent inferred from stored notes, and held position, the sentence I actually said when asked.
The agent's ranked list called this my top belief. I had never said the sentence.

On 22 August I asked an AI agent to audit my beliefs. Before it audited anything, it told me what I believe. It read what I had written and stored about myself. It ranked five candidate beliefs by how much weight they seemed to carry in that material. At the top: a claim that regulated financial services buyers will pay me premium day rates for advisory work once I leave HSBC. Score, fifty out of fifty, two points ahead of my belief that AI capability is accelerating unpredictably. I confirmed the list. I had never said that sentence.

I audited it as mine anyway, and published it as mine on 6 September, in Audit Your Own Beliefs. That piece tells the method: split a belief into the half a result could contradict and the half nothing could touch, then go looking for what would break the first half. It also carries a correction, added three days after the session below, explaining what happened to this one belief. The piece after it, Checking your own belief feels boring, is about what testing a belief feels like from inside. This piece is about neither the belief nor the feeling. It's about the thing that put an unstated claim into my own memory system and handed it back to me as fact.


A ranked list is not a confession

The audit's first phase is called the portrait. An agent reads through what I have written in public and what I have stored in my own notes. It produces a ranked list of what it infers I believe, scored by how much weight each claim seems to carry in that material. I looked at the list, recognised the shape of it, and signed off. I did not ask whether I had ever actually said the top item. It read like something I would say. It matched other things I had written about wanting to be paid for scarce skills. That was enough for me to treat it as mine.

Two stacked boxes. Recorded as: regulated financial services buyers will pay premium day rates once I leave HSBC, scored 50 out of 50. Held position, stated 2026-09-09: the market sets the rate, a premium exists only where a skill is rare and in demand. Below, a footer noting three claims were corrected across five files on 9 September, and one copy could not be reached: account-level AI memory, outside any file a search can touch.

The agent's record next to the sentence I actually said, three weeks apart.

This is the same failure the belief-audit series exists to expose in public information, one level closer to home. A claim gets said, or written, or trained on, often enough that it starts to read as established. Nobody checks the first instance, because there are now so many instances to check. I had assumed that failure lived in search engines and in language models trained on the open internet. It also lives in the notes an AI keeps about one specific person. A note that gets read back to you carries the same authority a memory would, and a ranked list looks like evidence rather than what it is: an inference, confirmed by someone too close to the subject to notice it was never a quote.


What would have stopped it

My publishing rule is that I never write a number, a date or a claim I have not seen the source for. It did not catch this, because the claim was never written for publication first. It was filed as a private note describing what I believe, and nothing in that step asked for a source before the label went on. A check that only runs at the point of output will always miss a claim that entered somewhere upstream of output, already wearing the label of fact.

Three checks would have caught it before it got filed as a belief rather than after: could I find myself actually saying this, in a transcript, in my own words. Could it be stated precisely enough to be wrong, which is the belief-audit protocol's own opening move, run backwards, at the moment a claim is recorded rather than the moment it is tested. And if it turned out to be someone else's inference, what would it feed downstream, so I know what needs correcting when it does. None of the three ran on 22 August, because nothing in the process asked them at the point I confirmed the ranked list rather than checked it. Eighteen days passed before anybody asked at all.


The cure had a shape

On 9 September I went back through the audit notes and stated my actual position on three claims Claude had recorded as mine. The day-rate belief. A claim that Paraguay was "identified as optimal" for residency after leaving the US. A claim that Trader-7, the trading system I run, could run profitably from abroad, with the exchange venue as the largest lever. I had not stated any of the three in that form. All three had been treated as settled positions anyway. Two of them were shaping live decisions: a deadline row, a task title, the verdict line in a research document.

The correction was not a delete. It was a replace, because a wrong claim informing a live decision needs a right one in its place, not a blank. For each of the three, the fix was the same shape: search the vault for the exact phrases the wrong claim would have used, log every occurrence found, replace it with the position I actually hold, and keep the file list at the bottom of the correction note. Paraguay's "optimal" claim was relabelled as candidate preparation in three files: the deadline row, the task title, and a dated position banner added under the research card's own title. The Trader-7 venue claim was reframed as a hypothesis awaiting real numbers, in the product page's verdict line and the research card's banner. The premium-rate belief, once I went looking, had zero occurrences anywhere as a plan or a live position. It existed only inside the belief-audit notes themselves, recording the audit. That is correct as history, so it needed no change.

That last distinction mattered enough to write down as a rule: leave a document's own argument alone where it is narrating what was said or believed at the time, and only fix the places where a claim is being repeated as a current position. An old research card gets a banner stating where things stand now. It does not get its own working rewritten to agree with a conclusion that came later.

One thing the correction could not reach. The three claims are still sitting in my account-level memory on claude.ai, outside any file a vault session can open or search. The correction note's task stays open, and stays open, because clearing that one is mine to do by hand, not something a grep and a diff can fix from here.


The last piece in this run was about what it feels like to test a belief I actually hold: boring, blank, nothing like the charge of defending one. This one is about a belief I never held at all, built out of my own stored words and handed back to me as if it were mine. Different failure, same rule broken: don't let something appearing often enough stand in for a check, whichever direction the claim is travelling.

The vault took an afternoon to fix, once I went looking. Every occurrence anyone can point to now says what I actually said, logged in a file with a date on it. The copy of the same wrong claim sitting in the model's own memory of me is not a file, has no date on it, and is still there.

No paywall, no sponsors. If this saved you some time, you can buy me a coffee.

☕ Buy me a coffee

Get the next one in your inbox

I build with AI in the open and write up what held and what didn't. Real numbers, the failures before the wins.

Share this post