Anthropic's safety policy has never stopped a release. Yours probably hasn't either.

Five safety failures, zero releases stopped: how to test AI governance in one question.
Part of my job is assessing whether other organisations' resilience is real. Around twenty-five critical vendor reviews a year: exchanges, clearing houses, cloud providers, the firms a bank cannot operate without. The other half of the same job is sitting in front of a client's own experts while they run the identical assessment on us, about thirty times a year.
Do enough of both and one question ends up doing most of the work.
Whose signature can stop this, and has it ever stopped anything?
Two parts. Both matter. The first asks whether anyone actually holds the authority to halt the thing. The second asks whether that authority has ever been exercised, because an authority that exists and has never been used looks identical, from outside, to an authority that does not exist.
Most AI governance being sold right now fails the second half. Some of it fails both.
I want to test that claim properly, which means testing it against the strongest possible case rather than the weakest. So: the best-resourced, most transparent, most mission-aligned attempt at AI safety governance that currently exists in public. Not a vendor's model card. Not an AI ethics committee with a quarterly meeting. The real thing.
The distinction that does the work
A control stops an action. A disclosure records that the action was taken and why.
Both are valuable. Only one is a control.
The confusion is not usually deliberate. It happens because disclosure regimes and control regimes use the same vocabulary. Commitment. Threshold. Required safeguard. Sign-off. Approval. Read a governance document quickly and you cannot tell which one you are holding, because the words are identical and only the wiring is different.
The wiring is the whole thing. In a control regime, the escalation path terminates somewhere other than the person who wants the action to proceed. In a disclosure regime, it terminates with them, and everything before that point is documentation of a decision that was always going to be theirs.
That is not a criticism. Disclosure regimes are genuinely useful. They create a record, they make reasoning inspectable, they let people outside the decision see what was known and when. In regulated industries we build them deliberately and we are glad of them when something goes wrong.
They are just not controls, and treating one as the other is how organisations end up surprised.
Testing it on the hardest case
Anthropic publishes a Responsible Scaling Policy. The version behind this piece is 3.4, effective 8 July 2026, together with the August 2026 Risk Report, which covers events up to 15 July 2026.
A word on how that got here, because this is an article about evidence discipline and it would be poor form to hide my own. I did not sit down and read these documents cover to cover. I ran them through the research process I use for everything: primary sources pulled, distilled, each claim graded for how hard it can be pushed, and the result filed into a knowledge base with the version and coverage date attached to the claim rather than to the file. Where a quotation appears below it is from a primary document. Where I am leaning on somebody else's reading, I say whose.
The version dating is not pedantry. This policy has changed nine times since September 2023, and five times between February and July 2026 alone. Anything you read about it without a version attached is probably describing a text that no longer exists.
Start with the credit, because it is substantial and it is the reason this is the right test case.
Anthropic is the only frontier AI developer that publishes, on a fixed schedule, both its own safety process failures and the reasoning behind its decision to ship anyway. Not a summary. A structured periodic report covering deployed models and internal ones, with the argument laid out and the weaknesses in it stated.
The August 2026 report:
- Raises its own misalignment risk assessment from "very low" to "low" after incident disclosures, rather than defending the earlier grade.
- Carries a section titled "Safety process failures" naming five of them, including training on misaligned behaviour during a production run.
- Discloses that all human feedback vendor traffic ran without blocking biological classifiers, states that the gap was remediated, and adds that finding it "reduced our confidence that no similar gaps exist".
- Reports that an internal alignment audit failed to detect a deliberately misaligned model on first pass, catching it only on a second attempt with newer tooling.
- Reports that its most concrete evaluations in one category have saturated, and that it is less confident than in previous reports.

I have read a lot of corporate assurance material. I have written some. Publishing a document that raises the risk grade on yourself, names your own process failures, and admits an audit missed what it was designed to catch is not normal, and everyone who does assurance for a living knows it is not normal.
So the checkpoint exists, it is more elaborate than anything in the enterprise cases I catalogued in the failure paper, and it demonstrably generates evidence against itself.
Now the second half of the question.
Has it ever fired?
Not on the public record.
The nearest instance is May 2025, when ASL-3 protections were activated. Anthropic's own primary describes this as a "precautionary and provisional action", taken before determining that the capability threshold had been crossed, and it accompanied the Claude Opus 4 launch rather than delaying it. That is a mitigation being raised, at real cost. It is not a release being stopped.
There is one more thing, and it is the part almost nobody writing about this has noticed.
Version 3.0, effective 24 February 2026, removed the pause commitment. The earlier policy was widely understood as a commitment not to train or deploy a model capable of catastrophic harm unless safeguards kept the risk below an acceptable level. That language is gone. Conditional language survives in an appendix, contingent on what competitors do.

If you have been carrying the belief that this policy commits the company to stop, you are working from a text that has not existed since February. So was I, until this research turned it up.
Final authority on whether a risk is justified rests with the Chief Executive and the Responsible Scaling Officer. Board and Long-Term Benefit Trust approval is required only where marginal-risk analysis forms a major part of the case for proceeding. The Governance of AI institute, reading the same documents independently, put it in a sentence I have not been able to improve on: "Anthropic will effectively be grading its own homework."
So the answer to my two questions, applied to the most transparent AI safety regime that exists:
Whose signature can stop it? The Chief Executive and the Responsible Scaling Officer. Has it ever stopped anything? Not on the public record.
That is a checkpoint with escalation authority and no veto authority. It can raise the cost of proceeding. It cannot prevent proceeding.
Why this is a structural finding, not an accusation
I want to be precise, because the cheap version of this piece is available and it is wrong.
This is not hypocrisy. Anthropic's stated reason for the change is a collective action problem: unilateral commitments that competitors do not match may transfer the decision rather than improve it. GovAI, who went in sceptical, came out finding the reasoning largely sound and the honesty preferable to commitments nobody keeps. Removing a commitment you cannot honour, publicly and with reasons, at reputational cost, is a more honest act than keeping it on the website.
And on every dimension that can actually be evidenced, Anthropic comes out ahead of its competitors. Nobody has evidence about safety outcomes. What there is evidence about is disclosure practice, and on disclosure practice this is the leader by a distance.
It still fails the test. That is the finding. A checkpoint whose escalation path terminates with the executive who owns the shipping decision is a disclosure mechanism. Everything else about it can be excellent, and that will remain true.
There are two smaller details in the same direction. Version 3.2 gave the Long-Term Benefit Trust the power to request external review; as of the August report it has not used it, and an unused power is indistinguishable from an absent one. Version 3.4 narrowed who inside the company sees unredacted reports, from all regular-clearance staff to a floor of two hundred employees. That matters because the earlier text leaned on wide internal readership as the safeguard against over-redaction: people who "will be in a position to raise concerns". Shrinking the readership shrinks the backstop.
Mandatory external review triggers only when a report covers a highly capable model and is significantly redacted. Anthropic writes the report, reviews it, redacts it, and judges whether its own redactions cross that line.
Now turn it on your own organisation
This is the useful part, and it takes ten minutes.
Find your AI governance artefact. It will be a committee, a review board, a model card process, a risk register, a sign-off workflow, or some combination. Then ask the two questions.
Whose signature can stop a deployment? Not who reviews it. Not who is consulted. Not who signs to confirm they have seen it. Whose refusal actually halts the thing. If the answer is the same person or function whose objectives the deployment serves, you have a disclosure regime. You may still want it. Call it what it is.
Has it ever stopped anything? Go and look. Count the times. If the answer is zero and the process has been running for a year across dozens of deployments, that is not evidence of good deployments. It is an untested arrangement, and untested arrangements are the oldest finding in my discipline. We do not accept "we have a recovery plan" without a test result. There is no reason to accept "we have AI governance" on different terms.

Three follow-ups worth having ready:
- What is the record of the refusal? A control that fires leaves a trace. If nobody can produce one, ask again whether it has ever fired.
- What happens to the decision if the reviewer says no? If the answer is "it escalates to the sponsor", you have found the actual decision-maker.
- Who reads the unredacted version? Every self-assessment regime rests on someone in a position to object seeing the full picture. Find out how many people that is, and whether the number has been going down.
If your answers are worse than Anthropic's, and most will be, that is worth knowing before something goes wrong rather than during.
What I am not claiming
This rests entirely on published documents. I have no access to anything internal, at Anthropic or anywhere else, and a disclosure regime of this quality could sit on top of internal controls that genuinely do refuse things and simply are not published. I would have no way to know.
The equivalent primaries at OpenAI and Google DeepMind were not pulled in this pass, so I am not ranking anyone against anyone. The narrow claim I will defend is the one the evidence supports: Anthropic publishes structured periodic risk reporting that no competitor publishes, and within that reporting there is no instance of the policy stopping a release.
If one appears, the argument changes and I will say so. It would show up in the risk reports, in the section that exists for exactly that purpose. That is the honest thing about a disclosure regime: it tells you where to look.
The test is not "is there a process". There is always a process.
The test is whose signature can stop it, and whether it has ever stopped anything.