Skip to main content

The credibility test we run backwards: ask what the warning cost

Jamie Watters

Operational resilience and AI delivery practitioner. Technology since 1985.

Published: 15 September 20266 min read
#ai-safety#credibility#evidence
Two statements side by side on a dark card. One cost unvested equity and a career and was dismissed. One cost nothing and was believed.
Same week, same subject, opposite receptions.

In the same week, two AI researchers said roughly the same frightening thing in public.

One of them had just quit his job and given up his unvested equity to say it. The other still had his job, and his chief executive had been saying a version of it for two years.

The panel that discussed them believed the second one.

Jacob Coxon spent about three years on pretraining research, first at OpenAI and then at Anthropic. He told TIME he walked away on 8 September. The next day he posted that neither company is acting responsibly, that they are "racing straight to self-improving superintelligence and gambling with our lives". He resigned before his Anthropic equity vested. He has left AI research, not just the employer.

Evan Hubinger leads Alignment Science at Anthropic. The same day, he replied on X: "Jacob is correct here, we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."

Peter Diamandis's Moonshots panel covered both in one episode. Coxon's resignation got "blaze of glory rage quit", "virtue signaling", don't give it much credibility. Hubinger's ten per cent was read out and taken at face value. There is no point in the episode transcript where anyone asks why the bar moved between the two.


What each sentence cost

Coxon paid money he had earned and not yet been given, plus his position in the field. That is a receipt. You can argue about whether he is right, but not about whether he meant it.

Hubinger's statement has no findable cost at all. He was in the role the day before and the day after. His chief executive, Dario Amodei, has been giving a public range of ten to twenty-five per cent since 2024. Amodei told Axios in September 2025 there was "a 25% chance that things go really, really badly". Hubinger had already published an essay arguing the industry should slow down. His number sits inside his employer's existing public position and continues his own.

That is not an accusation. It is an accounting. He may well be right. He simply didn't have to pay anything to say it, and the man the panel dismissed did.

Two cards compared. Jacob Coxon paid unvested equity and his position in the field, and the panel called it virtue signaling. Evan Hubinger paid nothing findable, with his chief executive already public on ten to twenty-five per cent, and the panel called him credible.


The test we think we're running

Ask people how they judge a claim and most will tell you they follow the incentives. Then watch which way they follow them.

A loud exit reads as drama. Staying reads as measured. The person who quit is assumed to be performing, and the person who stayed is assumed to be brave within the constraints. It feels like the sophisticated read. It is the incentive check run backwards, because the exit came with a receipt and the insider's sentence came free.

The bit that matters is narrower than "did they leave", and getting it precise is what makes it useful at work. The cost that counts is what saying it cost them at the moment they said it.

Your colleague who tells the sponsor in the meeting that the date will slip is paying, in that meeting, in front of the person it embarrasses. The same colleague writing it in a leaving email three weeks later pays nothing. Same words, same person, same slip. Only one of them is evidence.

Coxon didn't quit and then complain. He quit in order to say it, and the equity went with the sentence.


What cost cannot buy

Here is where it stops, and this is the part both sides of the argument skip.

Cost proves sincerity. It does not prove accuracy.

Coxon giving up his equity means he believes what he said. It does not mean he is right. Sincerity bought at a price is exactly how a mistaken belief acquires a reputation for integrity, and the history of costly warnings that turned out to be wrong is long.

The reverse holds too. Hubinger's statement being free does not make it empty. He introduced mesa-optimisation in a 2019 paper, and most of the field now borrows its vocabulary from it. He co-authored Sleeper Agents in 2024, which showed that a model can be trained to behave deceptively in a way standard safety training does not remove. He spends his days building the tools that detect the failure he is warning about.

So cost sorts belief from posture. It does not sort true from false. Two different questions, and the episode collapsed them: it treated "is he sincere" as though it answered "is he right", and then applied even the sincerity test to only one of the two men.

As for the number itself, it has no method attached. No model, no reference class, no date on which it resolves, no observation that would move it. A greater-than-ten-per-cent chance of extinction inside a decade forbids nothing you could go and look for next year. Hubinger was careful, to be fair to him, and said "I personally think". That hedge survived into most of the coverage. It is his belief, honestly labelled as his belief, and it is not a measurement.


Three questions for any warning

Three questions for any warning. One, what did saying it cost them, at the moment they said it. Two, is the claim checkable or only believable. Three, what would change their mind. Cost sorts belief from posture; it does not sort true from false.

  1. What did saying this cost them, at the moment they said it? Not whether they are angry, not whether they left. What did the sentence cost.
  2. Is the claim checkable, or only believable? Can you name the thing that would show it to be false.
  3. What would change their mind? Ask them. If the answer is nothing, you have a belief, not a forecast.

Run the two statements through it. Coxon's passes the first and says nothing about the other two. Hubinger's fails the second and third, and I suspect he would say so himself.

That is the useful result. Not "believe the quitter" and not "it's all doom marketing", but a way of telling what you have got in front of you before you decide what to do about it.

Same week, same subject, two men. One paid for his sentence and was called a virtue signaller. The other paid nothing and was believed. Nothing in the transcript shows anyone on the panel noticing which way round that was.


Sources

  • Evan Hubinger (@EvanHub), X, 9 September 2026. The quote is reproduced identically by Dexerto, Wall St Engine and BigGo Finance, each linking the original post.
  • Jacob Coxon (@hilbertspaess), X, 9 September 2026.
  • TIME, "He Helped Build Powerful AI at OpenAI and Anthropic. Now He's Afraid It Could Kill Us", September 2026. The interview is the source for the 8 September departure, the three years, the forfeited unvested equity and his leaving the field.
  • Axios, "Amodei on AI: 'There's a 25% chance that things go really, really badly'", 17 September 2025.
  • Hubinger and colleagues, "Risks from Learned Optimization in Advanced Machine Learning Systems", 2019, and "Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training", 2024.
  • Moonshots with Peter Diamandis, "Three Lab Warnings in Five Days", September 2026. The panel quotes come from the full episode transcript, not from a clip.

No paywall, no sponsors. If this saved you some time, you can buy me a coffee.

☕ Buy me a coffee

Get the next one in your inbox

I build with AI in the open and write up what held and what didn't. Real numbers, the failures before the wins.

Share this post