Anthropic's alignment lead put extinction above 10%: ask where the plan is, not whether the number is right
Jamie Watters
Operational resilience and AI delivery practitioner. Technology since 1985.

On 9 September, Evan Hubinger, who leads Alignment Science at Anthropic, wrote three sentences on X. Everyone quoted the middle one.
"I personally think it is >10% within the next decade." Greater than ten per cent, personally, that AI could kill all humans. The internet did what it does with a number: argued about it. Too high, too low, his own chief executive told Axios twenty-five a year ago, a podcast host says his was fifty and is now twenty.
Here is the third sentence, which travelled much less well: "I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."
That is not a footnote to the ten per cent. In the job I do, that is the whole entry.
What a number is supposed to arrive with
I have run business continuity and incident management for a global bank's markets business in the Americas since 2015. A large part of that job is reading risk registers and asking rude questions about them.
First, the words, because the piece falls over without them. Hubinger gave two different things. The severity is extinction: total loss, and nobody needs to estimate that. The ten per cent is a likelihood, and a likelihood is the softest thing on the page.
You cannot put a likelihood in a register on its own. It is not that it would be frowned upon. It is that the row is not a row yet. Here are the six things I want sitting next to it before I will sign anything. This is my own checklist rather than a standard: ISO 31000 prescribes no column count, and real registers run anywhere from seven fields to fourteen.
- How the number was derived. A model, a reference class, a scenario walk-through, a panel of named people using a stated method. Something that a second person could repeat and get a similar answer. It should also say whether it is the likelihood before the existing safeguards or the likelihood left after them, because those are different numbers and people quote them interchangeably.
- What the event actually is. Defined tightly enough that you would know if it had happened, or had started happening.
- Impact and horizon. How bad, over what period, measured in something.
- A named owner. One person, not a function. The person who gets asked.
- The controls already in place, and evidence they work. Not a list of intentions. Test results.
- A treatment decision and a review date. Accept, reduce, transfer or avoid. Written down, dated, and signed by someone with the authority to carry the consequence.
Count what the pdoom conversation actually supplies. The likelihood, stated, with no method behind it. The event, loosely: AI kills all humans. The horizon, a decade. Then three blanks: no derivation, no owner of what is left after the safeguards, no controls with evidence they work. And in the seventh cell, a public statement from the person best placed to fill it that the plan does not exist.
Three cells carrying something, three blank, and the seventh filled in with an absence. None of the three that carry anything tells you what is being done.
Item two is funny for about a second. The event is defined tightly enough: everyone dies. It is the verification that breaks. A register works because somebody comes back on the review date and checks whether the thing happened, and here the defining feature of the event is that there is nobody left to tick the box. Which is worth saying plainly, because it rules out the position most people are actually taking, which is that we will know more as it develops.

The number is not useless. It is insufficient.
I have been through US regulatory audits and reviews with the OCC, the NFA and the Federal Reserve. Not one of them ever opened with "we would like to challenge your likelihood rating".
They asked for three things. Show me the plan. Show me who owns it. Show me the evidence you tested it.
That is not because the number does not matter, and I want to be careful here, because the handbook those examiners work from has a section called "Likelihood and Impact" telling them to check that the risk assessment addresses both. Likelihood decides where the examination goes and how much of it you get. What it does not do is settle the meeting: two competent people give different numbers for the same risk and neither can prove the other wrong, so the weight lands on the parts that can be checked.
The number gets you the meeting. It cannot be the only thing you brought to it.
A firm that said "there is a real chance this brings us down, we have not worked out what to do about it, we are carrying on at pace" would not be praised for candour. It would expect a finding, and the conversation would turn to what stops until there is a plan.
I am not claiming the labs are in that position. No supervisor is examining them, and that is rather the point. There is a whole discipline whose job is to stop the jump from "here is an estimate" to "here is a policy debate", and I have not seen a single piece of coverage apply it here. The discourse treats the register as an optional formality. The register is the mechanism.
The strongest argument against all of this
A risk register governs risks inside one organisation's boundary, and that is exactly what this risk is not.
Suppose Anthropic did the full entry tomorrow: method, owner, controls, treatment, review date, signed. Suppose the treatment was avoid, and they stopped. The likelihood barely moves, because the competitors do not stop, the open weights are already out, and no state is waiting for Anthropic's sign-off. In my trade that has a name and a shelf: it is a systemic risk, the kind that gets logged as macroeconomic or geopolitical, where the honest corporate treatment is acceptance plus advocacy, because the firm has no unilateral lever. Demanding an internal plan for an uncoordinated global race is a category error.
That objection is right about the treatment and wrong about everything else. It kills "avoid" as a solo move and it kills nothing else. Method, event definition, owner of the residual, controls with evidence, and a signed statement of what is being accepted: all five survive the risk being industry-wide, and all five are missing. A systemic risk still gets a row. What it does not get is the pretence that one firm's row makes it go away.
A firm that writes "we accept this, here is why, here is who signed it" has done something this conversation has not. It has filled the cell in.
The part where my discipline does not fit, said out loud
If I stopped there I would be selling you something.
Start with the thing my trade is actually good at, because it is not what people think. We do not plan for risks. There is an infinite list of ways to lose a building and only one outcome that matters, which is that the building is gone. So the plans are written against outcomes rather than causes: loss of premises, loss of people, loss of a supplier, loss of systems and data. Fire, flood, a burst main, a police cordon, it makes no difference by Tuesday. Four plans cover the infinite list, and that is the trick that makes any of this tractable.
It is also why arguing about the probability of a specific cause is not where the work is. We gave that up a long time ago and went to work on the handful of outcomes instead.
Run this risk through the trick and it fails. The outcome is loss of everything, permanently, including whoever would have run the plan. You cannot rehearse a failure mode that has not been described in enough detail to stage, and you cannot write a recovery time objective for an unrecoverable event.
A disaster recovery manager who worked for me, Ray Trapnell, used to state the boundary plainly: if a nuclear bomb goes off, I'm heading home. He was not being flippant. He was making the professional point that somewhere past a certain scale the honest answer is that this is not a continuity problem, and that saying so out loud is better practice than writing a plan nobody could execute. Half of what I do for a living does not transfer to this, and anyone telling you it does is reaching.
What transfers is the other half, and it is the unglamorous half: the row, not the drill. A stated method. A defined event. An owner. Known controls with evidence. A treatment decision with a date on it. None of that requires you to understand the mechanism. It requires you to be honest about what you have got.
And the disanalogy points somewhere useful rather than shutting the question down. Lose rehearsal and recovery and you lose the corrective controls, the ones that act after the event. What is left is prevention and detection: stopping it, and seeing it coming early enough to matter. That is a thinner set than any resilience professional would like, and it puts almost all the weight on the two treatments that act before the fact, reduce and avoid. Accept stays on the list, because it always does, but accepting a catastrophic risk is a decision someone signs rather than a thing that happens by default.
Three days after the exchange, Hubinger's own chief executive published an essay arguing the industry should pace the frontier, and Hubinger endorsed it. That is the harder conversation, and it is downstream of the register rather than of the number.
The test, which travels well beyond AI

Next time somebody hands you a percentage, ask four things. It does not matter what the percentage is about: extinction, a project someone is confident will land in March, a security warning, a health scare.
- How did you get that number? Method, not confidence.
- What exactly is the event? If it cannot be described tightly, the percentage is decoration.
- Who owns it? A name.
- What is the plan, and when was it last tested? This is the one that does the work.
Most of the time you will get an answer to the first and nothing on the fourth. That tells you what you are holding: a belief, sincerely held, with no machinery behind it. The number is still worth having. It tells you how worried to be and how fast to move. It just does not tell you what to do, and it is the cheapest part of the row to produce.
Sometimes a belief with no machinery is all anyone has, and saying so plainly is a perfectly respectable position. Pretending the number is the substance is not.
I wrote about the other half of this story yesterday: a researcher who quit and forfeited unvested equity to say the same thing was dismissed as a virtue signaller, while Hubinger, who paid nothing, was believed. That piece was about who to believe. This one is about what to ask them next.
Hubinger gave us both halves in the same three sentences. The argument went to the number. The plan is the part that is missing, and he is the one who told us so.
Sources
- Evan Hubinger (@EvanHub), X, 9 September 2026. Quoted from the original post; the same text is reproduced by Dexerto and BigGo Finance, each linking it.
- Axios, "Amodei on AI: 'There's a 25% chance that things go really, really badly'", 17 September 2025.
- Dario Amodei, "We Must Pace the Frontier", published 12 September 2026, three days after the exchange above, and endorsed by Hubinger the same day. Corrected 15 September after peer review: an earlier version of this piece gave the essay to Hubinger and placed it before the exchange. Both were wrong, and the error came from my own research note, which is now fixed at source.
- FFIEC Business Continuity Management booklet, section III.B.2, "Likelihood and Impact", which is the handbook the examiners in this piece work from.
- ISO 31000:2018, for the separation of likelihood from consequence and for the four treatment options.
- Moonshots with Peter Diamandis, "Three Lab Warnings in Five Days", September 2026, for the host's own figures, read from the full episode transcript rather than a clip.
- My own record for the regulatory examinations referred to: US audits and reviews with the OCC, the NFA and the Federal Reserve, completed with no business continuity issues raised.
Revised 15 September after five independent reviews. Four of them found the same wrongly attributed essay, and the strongest objection in the piece, the one about a systemic risk having no owner inside any one firm, is in it because two reviewers put it there.