Skip to main content

AI arms race game theory: the game we're in turns on an unmeasured belief

Jamie Watters

Operational resilience and AI delivery practitioner. Technology since 1985.

Published: 16 September 202612 min read
#ai-safety#game-theory#ai-governance#agi#risk
Social card reading: a prisoner's dilemma in one world, a coordination problem in the other, and what decides which is a shared belief the model assumes rather than measures.
RAND's two worlds. In the second one restraint is only as stable as racing, which makes it a coordination problem rather than a solution.

Who this is for. Anyone who has heard, or said, that there's no choice but to build advanced AI as fast as possible because someone worse will otherwise get there first. By the end you'll have four questions to put to that argument.

Skip it if you want a prediction about how the race ends. This is about the structure of the argument, not the outcome.

On the length. Twelve minutes, and the four questions at the end work on their own.


Almost everyone arguing about AI risk is playing the same game, on three premises nobody says out loud.

One: artificial superintelligence is the ultimate power. Two: it must not fall into the wrong hands. Three: we are the good guys.

Take those as given and a strategy drops out with no further thinking required. Build it first. Build it fastest. Treat any pause as handing the win to someone worse. You hear the conclusion constantly, usually as "we have no choice", and you almost never hear the premises, because saying them out loud invites the obvious question about the third one.

I wanted to know whether that reasoning survives contact with the people who model this formally. It doesn't, and the way it fails is more useful than a debunk.

The label is a specific claim, not a mood

The word people reach for is prisoner's dilemma. It usually means "bad situation nobody can escape", which is not what it says.

In a one-shot prisoner's dilemma, defection is the dominant strategy. Whatever the other side does, you're better off racing. Mutual restraint would leave you both better off than mutual racing, and it still isn't an equilibrium, so nothing holds it in place.

A stag hunt is different in the way that matters. Mutual restraint and mutual racing are both equilibria. Neither side holds a dominant strategy, and the payoffs leave both outcomes available, so what each side expects the other to do decides which one you land on.

That difference changes what there is to do. A one-shot prisoner's dilemma needs something that alters the payoffs from outside: an enforcer, a treaty with teeth, a cost for defecting. A coordination problem needs a way to believe the other side, which is cheaper, and which is why it matters whether you have guessed the game right.

Which one is it? The honest answer is conditional

RAND published a model built to test exactly this, in December 2025, by Lisa Abraham, Joshua Kavner and Alvin Moon. Its title is "A Prisoner's Dilemma in the Race to Artificial General Intelligence", and the paper is more careful than the title.

It models two countries choosing whether to accelerate, and it gets two worlds.

In the first, the report's own words: "If both the United States and China have the same shared assessment that the benefits of being the first to achieve AGI outweigh the risks, they are effectively locked in a prisoner's dilemma." Acceleration is the only equilibrium and both do it.

In the second, when the shared assessment flips, the report says mutual restraint becomes "as stable as mutual acceleration", and that this turns the problem "from a dilemma, in which defection is always tempting, to a coordination challenge".

Read that second one carefully, because it isn't a rescue. Restraint being as stable as racing means both are still available. You get a coordination problem, which you can still lose.

Now the assumption holding it up, stated plainly in the paper: "All countries have the same information and correspondingly have the same perceptions about the benefits of first-mover advantage and the risks of accelerated development."

That is an assumption, not a measurement. And when I went looking for anyone who had measured it, a survey of decision-makers, an elicitation protocol, an attempt to infer the belief from budgets or compute allocation, I didn't find one. I searched the four places it would most likely appear. That's "I looked and didn't find it", not "it doesn't exist".

Which is enough for the point. The argument that we have no choice but to race is a bet on an unmeasured belief, delivered in the tone of arithmetic.

Two panels. Left, World 1: both sides share the assessment that being first outweighs the risk, so the game is a prisoner's dilemma and acceleration is the only equilibrium. Right, World 2: both sides share the assessment that catastrophic risk outweighs being first, and mutual restraint becomes as stable as mutual acceleration, which is a coordination problem rather than a solution. Underneath: the model assumes both sides hold the same assessment, and I found no measurement of it.

Source: Abraham, Kavner and Moon, "A Prisoner's Dilemma in the Race to Artificial General Intelligence", RAND RR-A4245-1, December 2025.

Several people got here at once, and I'm not going to pretend otherwise

This is the part where I'd normally present the observation as a find. It isn't one.

Vaughn Papenhausen made the same move on 4 August 2026, arguing the game shifts between prisoner's dilemma and stag hunt at a threshold set by the ratio of the winning payoff to the doom payoff. A 2026 preprint on a superintelligence moratorium says in its own limitations that "it might be interesting to introduce each player's, potentially asymmetric belief of the cost", which is a polite way of noting that their model treats that cost as a shared given.

And Emile Naude posted a working paper on 15 September 2026, the day before I started, which models states and laboratories together and disagrees with the framing above. In his baseline, "laboratory responses rule out a Stag Hunt", and as catastrophe intensity rises the game "passes from a Prisoner's Dilemma through Chicken to dominant restraint". Both phrases describe that baseline rather than every version of his model, and it's a single-author working paper with no stated peer-review status, which discloses substantial language-model assistance in the drafting. Take it as a recent argument rather than as authority. But two formal models reaching different answers about where restraint lives is the actual state of this question, and the honest version of this piece is that several of us landed in the same place at once rather than that I found something.

Premise three has a name

"We are the good guys" can't be true for every player at once. Every player acts as though it is.

The temptation is to call that hypocrisy. Psychology has a better answer. Robert Robinson, Dacher Keltner, Andrew Ward and Lee Ross studied it in 1995 under the name naive realism: people tend to feel that they themselves reasoned from the available evidence to a sensible conclusion, while those who disagree did the opposite. Their finding was measured on student samples arguing about abortion and about politics, so treat it as a demonstrated human tendency rather than a law about states. What it predicts is the part that travels: partisans underestimate how much common ground is there, because they have misread why the other side disagrees.

Robert Jervis had described the same shape in international relations in 1976. A state that means no harm knows it means no harm, and assumes everyone else can see that too. The failure isn't in the first half of that sentence.

So premise three doesn't require anyone to be lying. It only requires each side to read its own reasoning more generously than it reads everyone else's, which is ordinary, and which requires no one to be caught out.

Premise one is assumed, by the people who argue from it hardest

The sharpest version of the racing argument turns out to be an argument against racing. Katzke and Futerman's preprint makes the case that "the same assumptions that might motivate the US to race to develop ASI also imply that such a race is extremely dangerous", and concludes that superintelligence "presents a trust dilemma rather than a prisoner's dilemma".

The chain runs through one condition: that a leading system grants a decisive advantage to whoever holds it. Without a decisive advantage there's no discrete moment to race towards, and the fear organising everything loses its object.

They don't establish that condition, they say so, and it's deliberate. They grant the racing advocates' premises and follow them to their consequences, which is what makes it an internal critique rather than a rebuttal. Their own line is "we do not attempt to resolve these uncertainties". That's an honest way to argue. It also means the premise doing the most work in public debate is carried, in the strongest formal treatment of it I found, as something granted rather than shown.

Every intuitive fix behaves counter-intuitively in these models

If you'd asked me at the start what makes a race safer, I'd have said more transparency and better safety technology. Both are more complicated than that, and each result below belongs to its own model rather than being a general law.

Armstrong, Bostrom and Shulman modelled teams that can trade safety for speed, and found ignorance safer than either of their information cases: "the more teams know about each others' capabilities (and about their own), the more the danger increases." Their phrase is that we'd be better off not knowing.

Fudenberg and Koh, in July 2026, found the effect isn't even monotonic, because faster detection makes it cheaper to wait and see whether your rival stops before committing to stop yourself. In their model, "at intermediate trust, increasing transparency can first destroy the early-stopping equilibrium before restoring it as detection becomes fast enough to make stopping self-enforcing."

Stafford, Trager and Dafoe found that safety improvements can be spent rather than banked: "when competitors are not deploying the riskiest technologies, steps to make those technologies safer will be attenuated or reversed by risk compensation." They also find a laggard's behaviour depends on where it sits, and not in a straight line: "if competitors are highly adversarial and the laggard is closer to the leader's capability level, the laggard is willing to cut corners to gamble for advantage", while a laggard that actually catches up can lower the shared risk.

None of that says transparency is bad or safety work is wasted. It says the intuitive version of each is not what these models produce, which is a reason to ask what mechanism you're relying on rather than assume the obvious one.

What restraint has needed, historically

One pair of treaties makes the point better than a survey would.

The Chemical Weapons Convention has challenge inspections. In the OPCW's own words, states "have committed themselves to the principle of 'any time, anywhere' inspections with no right of refusal". A state can't simply refuse a properly made request. The Executive Council can stop one within twelve hours, but only by a three-quarter majority and only if it's frivolous, abusive or outside the convention, and managed access still has to give inspectors the greatest degree of access consistent with protecting unrelated secrets.

The Biological Weapons Convention has nothing like it. Its enforcement hook, Article VI, lets a state "lodge a complaint with the Security Council of the United Nations", which may then investigate. That route exists and has been used: Russia brought one in 2022 and it failed for want of nine votes. What it isn't is an inspection regime. The Soviet Union ran an offensive biological weapons programme for years after ratifying the treaty in 1975, and the treaty had no machinery of its own that could have found it.

Two conventions, comparable ambitions, and only one has a way to check. That's the question worth carrying to AI: not whether anyone has promised restraint, but whether anything could show it.

The obvious objection, and the one that's stronger

The obvious objection is that if the game turns on a shared belief, you should measure the belief.

James Fearon's 1995 work is the standard reason that's harder than it sounds. States "wish to obtain a favorable resolution of the issues", and that "can give them an incentive to exaggerate their true willingness or capability to fight". In his bargaining model, announcing your position doesn't settle it, because "regardless of B's true willingness to fight, B does best to make the announcement that leads to the smallest grab by A". That's a result inside a cheap-talk model rather than a proof that every survey fails, but the direction is the point: the parties whose beliefs decide the game have a reason to misstate them.

The stronger objection comes from the same paper, and I'd rather raise it than leave it sitting. Fearon's other mechanism is the commitment problem. Two states that agree catastrophe would be terrible can still end up fighting, because neither can credibly promise not to use an advantage once it has one. Apply that here and a shared assessment stops being sufficient: if you believe a leading system confers power that can't be handed back, racing can be rational even with the ratio agreed. That doesn't rescue "no choice", because it swaps one unmeasured belief for another about how decisive and how irreversible the advantage would be. It does mean the ratio isn't the only thing to ask about.

Four questions

Put these to anyone who tells you there's no choice, including yourself.

  1. Which game are you claiming this is? Does defection really dominate whatever the other side does, or is there a cooperative outcome you've assumed away? Only one of those needs an outside enforcer.
  2. What is your benefit-to-catastrophe assessment, and do you think the other side shares it? The models turn on this, and "obviously huge" is not an answer. If you can't say what the other side believes, you're guessing at which game you're in.
  3. What would have to be observable for restraint to be credible here? History says this is the hinge. If the answer is nothing, you've conceded that promises are all anyone has.
  4. What has actually happened, as opposed to been announced? Statements of intent are cheap by construction, which is Fearon's whole point. Anthropic said in September that it had signed an agreement with METR to investigate a set of incidents, which is a real thing with a name and a scope attached. The permanent embedded-evaluator team it also proposed is a commitment rather than an arrangement. Tracking that distinction is most of the work.

None of that tells you whether the race ends well. It tells you that the sentence doing most of the work in this argument, the one saying there's no choice, carries an unmeasured belief about two quantities, and that the people who model this formally have not agreed on which game we're in.


Sources

  • Lisa Abraham, Joshua Kavner and Alvin Moon, A Prisoner's Dilemma in the Race to Artificial General Intelligence, RAND RR-A4245-1, December 2025, for the two worlds, the quoted conditions and the shared-information assumption.
  • Stuart Armstrong, Nick Bostrom and Carl Shulman, Racing to the Precipice: a Model of Artificial Intelligence Development, Future of Humanity Institute Technical Report 2013-1, later in AI & Society, for the information result in that model.
  • Drew Fudenberg and Andrew Koh, Racing to Ruin, arXiv 2607.27638, July 2026, for the non-monotonic transparency result.
  • Eoghan Stafford, Robert Trager and Allan Dafoe, Safety Not Guaranteed, Centre for the Governance of AI, November 2022, for risk compensation and for the laggard result, including its non-monotonic shape.
  • Robert Robinson, Dacher Keltner, Andrew Ward and Lee Ross, Actual Versus Assumed Differences in Construal, Journal of Personality and Social Psychology, 1995, for naive realism, measured on student samples. Robert Jervis, Perception and Misperception in International Politics, 1976, for the same shape in international relations.
  • Corin Katzke and Gideon Futerman, The Manhattan Trap, arXiv preprint, for the internal critique and its stated assumptions.
  • James Fearon, Rationalist Explanations for War, 1995, for the incentive to misrepresent and for the commitment problem.
  • Elias Fernández Domingos and The Anh Han, arXiv 2607.26034, for the laboratory experiment on how people behave inside a race they have been told they are in.
  • Vaughn Papenhausen, The AI Race is Not a Prisoner's Dilemma, 4 August 2026; Emile Naude, Nested Races, working paper, 15 September 2026; and arXiv 2605.01297 on a superintelligence moratorium.
  • Chemical Weapons Convention Article IX and the OPCW's own description of challenge inspections; Biological Weapons Convention Article VI and the 2022 Security Council vote.
  • Dario Amodei, We Must Pace the Frontier, September 2026, and Anthropic's statement of 9 September 2026 on the METR agreement.

Corrected after a round of outside review. An earlier version said Anthropic had no named evaluator and no live agreement; it had announced an agreement with METR on 9 September, and question four now says so. It compared Armstrong and colleagues with a paper by Cimpeanu and colleagues as opposite findings on the number of players, which was a category error, because that paper varies network structure rather than player count; the comparison is gone. It described RAND's second world as one where restraint becomes stable, where the report says restraint becomes as stable as acceleration and calls the result a coordination challenge. It said the premise carrying the racing argument is one nobody argues for, when Katzke and Futerman grant it deliberately as part of an internal critique. It misnamed two authors. And it ran a compressed survey of six historical regimes with claims stronger than the sources supported; two treaties remain, because they were the two carrying the argument.

No paywall, no sponsors. If this saved you some time, you can buy me a coffee.

☕ Buy me a coffee

Get the next one in your inbox

I build with AI in the open and write up what held and what didn't. Real numbers, the failures before the wins.

Share this post