Open-Weight AI: The Prisoner's Dilemma You Can't Escape
The open-weight AI race is a Prisoner's Dilemma you can't escape — because a published model can never be recalled. Here's the game, tested against nuclear, bio and climate, and the one lever that ends it.
Executive summary
The most important decision in artificial intelligence right now is not being made in an alignment paper or a training run. It is being made in a licensing choice — the choice to publish a model's weights or keep them locked — and it is being made, over and over, in the direction of no takebacks. China is the loudest instance: its open-weight models now carry the majority of the world's AI usage, sit only months behind the closed frontier, and undercut Western prices by up to two orders of magnitude. But the China story is a symptom. The disease is a game — a specific, nameable game — that every serious player is now trapped inside, and that almost nobody in the race will name out loud.
The game is the Prisoner's Dilemma, in its harshest form. Each lab and each nation faces the same payoff logic: whatever the other side does, the individually rational move is to defect — to release capability and ship — because unilateral restraint just means you lose while the other side ships anyway. Sign the safety statement, then ship. Mutual defection is the equilibrium even though mutual restraint would leave everyone safer. That much is a textbook trap, and textbooks tell you how to escape it: play it over and over, punish defectors, and the shadow of the future pulls you back toward cooperation. The catch — the thing that makes this dilemma worse than the one in the textbook — is that you cannot un-publish a weight. Every defection is permanent and cumulative. That single feature collapses the repeated game back into a one-shot game and deletes the only escape route game theory offers.
So the honest question is not "is China reckless?" but "has any species of this game ever ended well?" History's answer is narrow and specific. Irreversible, many-player, hard-to-verify dilemmas get solved only after a catastrophe forces the political will, and only where the dangerous capability is scarce and countable enough to verify a deal. Nuclear weapons cleared that bar — barely, and only after the world stared into the abyss in 1962 — because you cannot email a bomb. Biology never cleared it, because a method is cheap to copy and impossible to inspect. AI weights are biology, not physics: a file, copyable forever, un-recallable. The one thing that could drag AI from the biological column toward the nuclear one is the single chokepoint that is still scarce and countable — computing hardware, the fissile material of AI. That is the whole ballgame. It is also the fragile part: if training ever stops needing a mountain of chips, the last chokepoint disappears and the game converges on the worst historical base rate we have. This report names the game, tests it against five real histories, plays it forward, and maps the actual ways out — not hand-waving optimism, and not doom.
Why this belongs in the story of the AI revolution
Every episode of this show circles one question: who ends up holding the most powerful technology humans have ever built, and on what terms. For two years that had a comfortable answer — the frontier lived inside a few American labs, and the lever was to keep it there. The open-weight surge broke that frame, and the instinct has been to argue about China: is it naive, is it reckless, is it state-directed. Those are the wrong altitude. The reason a company in Hangzhou and a company in San Francisco keep making the same irreversible choice is not national character. It is that they are playing the same game and reading the same payoffs, and the game has a name and a known mathematics. Name it, and you stop arguing about the players and start reasoning about the ending. That is the move this report makes — because the decision being made in that licensing choice is, quietly, one of the most consequential of the whole revolution, and it is being made without anyone treating it as the strategic problem it actually is.
Open weight is not open source — and the last row is a one-way door
Start with a distinction almost everyone collapses, because the collapse hides everything that follows. A neural network is, in the end, a giant pile of numbers — the weights that encode what the model learned. "Open weight" means a lab publishes that pile: download it, run it offline, fine-tune it, build on it, no permission required. What you do not get is the recipe — the training data, the cleaning and training code, the exact process. You get the cake, not the ingredients.
"Open source," in the strict sense the Open Source Initiative codified in 2024, is a higher bar: enough released that a skilled person could rebuild a substantially equivalent model — data, code, and weights, under open terms. By that measure almost nothing marketed as "open source" qualifies. Meta's Llama ships weights with usage restrictions and no data, so the OSI does not count it. DeepSeek releases open weights under a permissive license but never published the data. The words describe two different gifts: open source is "here is how to make one," which spreads knowledge; open weight is "here is a finished one," which spreads dependency. The race runs almost entirely on the second.
Hold onto that last row. A closed model is a service the maker can throttle, patch, or switch off. An open-weight model is a fact in the world that no one can undo. That is not a licensing footnote — it is the hinge the entire game turns on, and we will keep returning to it. Publishing a weight is walking through a door that only opens one way.
China now runs the world's AI — and openness is why
If you still picture Chinese AI as cheaper knock-offs, update the picture. The striking number of 2026 is not a benchmark — it is usage. On OpenRouter, the largest neutral marketplace routing traffic across every major model, Chinese open models went from under 2% of tokens a year ago to roughly 61% by mid-2026. In a single February week, Chinese models processed about 5.16 trillion tokens against 2.7 trillion for US models; four of the world's five most-used models were Chinese. When independent developers spending their own money route the majority of real workloads through Chinese open models, that is the market voting.
It feeds a quieter figure that should stop any US strategist cold: by the US-China Economic and Security Review Commission's March 2026 analysis, roughly 80% of American AI startups build on Chinese base models. And the models are not bad — Moonshot's Kimi K2.6 became the first open-weight model to beat a top-tier closed model on a demanding coding benchmark, the frontier-and-give-it-away moment we walked through in our Kimi K3 episode, number 33, a few weeks back. The UK's AI Security Institute, which has every incentive not to exaggerate danger, finds leading open-weight models now trail the closed frontier on dangerous cyber capability by only four to seven months, down from six to ten. And DeepSeek runs roughly 35 to 100 times cheaper per token than the closed flagships. Capability within months, price within cents: for the enormous middle of the market, the Chinese open option is simply the rational default.
Giving it away is the strategy, not the sacrifice
Handing a frontier model to the world looks like it cannot be a strategy. It is one of the oldest in technology, and it has a name: commoditize your complement. Every product has complements — things worth more when your product gets cheaper. If you sell razors, blades are the complement; if you sell chips, the software is the complement. Drive the price of the thing next to what you sell to zero, and your own layer captures the value.
The AI stack has three layers: models, the compute they run on, and the applications on top. American labs sell the model — the metered product their valuations rest on. So the most damaging move a challenger can make is not a slightly better model; it is to make the model itself free, draining profit and power out of the model layer and into the two layers beside it — the compute underneath and the applications on top. Those are precisely the layers China is positioned to own: it manufactures the world's mid-tier hardware and has the industrial base to deploy AI into the physical economy at scale. DeepSeek's Liang Wenfeng said the quiet part out loud — "moats created by closed source are temporary," even OpenAI's — and frames open weights as the way you build the developer gravity that turns a model into an ecosystem. This is not generosity. It is detonating the price of the one thing your competitor sells, then moving the game to a board where you have the better position. It is the coldest read available to the player who is behind on the frontier but ahead on manufacturing.
Is it state-directed? Yes and no, and the "no" matters. Beijing has positioned China as the world's provider of open AI, pushed diffusion through its "AI+" program, and in July 2026 convened 29 countries to found a new World AI Cooperation Organization to shape global standards. Open releases still carry the content controls the state requires — trained in before the weights ship — so "open" means open to build on, not open to say anything. But the labs are not opening models because a bureaucrat ordered it; they are opening them because, for a challenger, it is independently the rational move, and the interests converge. That convergence is more durable than an order — no one has to be coerced — but also more fragile in one specific way: if openness ever started arming rivals more than building dependency, the same alignment could reverse. Hold that thought.
They signed the safety letter, then shipped anyway
So do the people building these models understand the danger? On paper, completely. China's top AI scientists are in the room where the alarms are raised. The International Dialogues on AI Safety, which has produced some of the strongest joint statements on frontier risk — including proposed "red lines" development should not cross — is co-convened by Turing Award winner Andrew Yao, who has warned that "we have suddenly found a way to create a new species many times more powerful than we are." Zhang Hongjiang of the Beijing Academy of AI sat at the same table. In June 2025 China stood up its own safety body, drawing in the same senior academics. In December 2024, seventeen Chinese firms signed AI Safety Commitments mirroring the West's. The comprehension is not in doubt.
What thins out is the gap between the statement and the shipping. When an independent index scored frontier developers on whether they had real, published safety frameworks — concrete thresholds, red-teaming, evidence they were hunting for danger — the leading Chinese labs scored at the bottom, several with no public framework at all, and DeepSeek quietly dropped off the second round of commitments. That is worse than naivety in one exact sense: naivety is fixed by information, and this is a revealed preference under competition, which information does not fix.
Here is the turn that makes it hard: that revealed preference is not Chinese. Ask the same skeptical question of the American labs. Do they understand the danger? Better than anyone — they write the papers. Meta open-weighted Llama for the identical commoditize-the-complement reasons. The leading US labs race so hard their own safety researchers periodically resign warning that speed is winning — the pattern we traced in episode 29, on why AI's own safety warnings fuel the race. "Sign the safety statement, then ship anyway because the other guy will if you don't" is not a national characteristic. It is the logic of a race with a winner-take-most prize — which is the first clue that we are not looking at a country. We are looking at a game.
Name the game: the Prisoner's Dilemma
Give it its real name. Two players, each with two choices — cooperate (restrain: keep the most capable weights closed) or defect (release and ship). Four outcomes. If both restrain, both are safer and the frontier stays containable: the best collective result. If both defect, capability floods out uncontained: worse for everyone, but stable. And the two off-diagonal boxes are where the trap springs — if you restrain while the other defects, you carry all the cost of falling behind and get none of the safety, because the capability is out there anyway. So look down each column: whatever the other side does, your best individual move is to defect. Defection is the dominant strategy, and two rational players converge on the one outcome that is collectively second-worst. That is the Prisoner's Dilemma, and it is not a metaphor here — it is the literal structure of the choice.
This is why "just show restraint" is not advice — it is a request to play the dominated strategy. Unilateral restraint does not make the world safer; it makes you the loser in a world that is exactly as dangerous, because the other player's release already lowered the floor for everyone. The equilibrium is not a moral failing of the players. It is the mathematics of the box.
It is a security dilemma, and offense is winning
There is a sharper version of this trap in international relations, and it fits AI weights unusually well: the security dilemma. Robert Jervis's classic result is that even states with no aggressive intent end up racing and sometimes fighting, and that the intensity of the trap depends on two material facts — whether offense or defense has the advantage, and whether you can tell one from the other. When offense dominates and you cannot distinguish a defensive build-up from an offensive one, mistrust and arms-racing become rational for everyone, and war gets more likely.
Publishing capable weights is close to a pure offense-dominant move. A released model helps the attacker — the proliferator, the person seeking uplift for a cyber or bio operation — more than it helps the defender, because the defender was already protected by the model staying closed, and now has to defend against everyone who downloaded it. Worse, you cannot tell a "we open-sourced this to democratize research" release from a "we open-sourced this to arm the ecosystem" release; the artifact is identical. So the openness that looks like transparency reads, through Jervis's lens, as the single worst configuration: offense-dominant and indistinguishable. That is the theoretical reason the good-faith case for openness and the arms-race case for openness produce the exact same action — and why no one watching from outside can tell which one they are looking at.
The honest counter deserves its place, because openness is not pure loss. A published model can be inspected by thousands of outside researchers who would never see inside a closed API; red-teaming is democratized; no single company gets to be the unaccountable gatekeeper of the most powerful tool ever built. Those are real defensive benefits, and for today's models the serious argument is genuinely close. What is not close is the direction of the trade-off: the more capable the model, the more the irreversible offensive uplift outweighs the transparency gain, because a defender can patch a disclosed flaw but cannot un-teach a released capability. Openness helps the defender at a fixed level of capability; it helps the attacker permanently as capability climbs — and capability climbs every quarter.
Why this dilemma is worse: the door only opens one way
Now the part that separates this from every Prisoner's Dilemma in the textbook — and it is genuinely good news that we understand it, because it tells you exactly where to push. A one-shot Prisoner's Dilemma is a trap. A repeated one usually is not. When the same players face the choice again and again, cooperation can emerge on its own: you punish a defection by defecting back, reputations form, and the shadow of the future — the knowledge that today's betrayal invites tomorrow's — pulls rational, selfish players toward cooperating. This is the most celebrated escape hatch in game theory. Tit-for-tat won Robert Axelrod's famous tournaments precisely because it could retaliate and then forgive, resetting to cooperation once the other side did.
Every part of that escape depends on one assumption: that a defection can be answered and the game reset. Open weights break it. You cannot un-publish. Once the numbers are mirrored across the internet they are permanent, and the safety training inside them can be stripped out in minutes with a few dozen examples. So there is no punish-and-reset. Every defection is permanent and cumulative — it stacks onto every prior defection and never decays. That collapses the repeated game back into a one-shot game, and the one-shot game is the trap with no exit. The shadow of the future cannot discipline a move that casts a permanent shadow of its own.
And two more features stack the odds. This is not a two-player game — it is N-player: many labs, several nations, and a global open-source community, any one of whom can defect for everyone. And it is hard to verify: unlike a missile silo, a training run leaves little external trace, so even a side that wanted to cooperate cannot easily prove it did, or catch a partner who didn't. Irreversible, many-player, un-verifiable — that is the worst combination of properties a cooperation problem can have. Which raises the only question that matters: has humanity ever solved a game with those properties?
Test it against history: five precedents, one scorecard
We have run this experiment before, in other domains, and the record is unusually clear about what makes the difference. Five precedents are worth putting side by side — nuclear weapons, biological weapons, the climate, naval arms control, and the plain business version of the same move. Read them not as loose analogies but as data points on two axes: could the dangerous thing be verified, and could a mistake be reversed. Those two axes predict the outcome better than any amount of goodwill.
The rest of this section walks each row, because the disanalogies are where the real lesson lives.
Nuclear: the optimistic bound AI cannot reach
Nuclear non-proliferation is the success story people reach for, and it is a genuine partial success — the number of nuclear states grew far slower than mid-century forecasts feared. But look at why, because every reason is a feature AI lacks. First, it took a near-catastrophe to create the political will: the world cooperated on test bans and the Non-Proliferation Treaty in the years after the Cuban Missile Crisis, when both superpowers had looked directly at annihilation. Cooperation was not reasoned into existence; it was frightened into existence. Second, and decisively, the dangerous input is verifiable. A bomb needs fissile material — highly enriched uranium or plutonium — which is scarce, expensive, physically detectable, and countable. Satellites and seismic sensors can catch a test; IAEA inspectors can audit a stockpile; enrichment leaves a signature. The whole regime rests on the fact that you can verify a country's fissile material to a reasonable approximation.
Now the disanalogy, which is the entire point. You cannot email a nuclear weapon. Its danger is inseparable from a scarce physical substance, and that scarcity is what made the deal checkable and therefore keepable. An AI weight is the opposite: it is information, copyable at zero cost, impossible to count once released, invisible to any inspector. Nuclear is the best case for governing a catastrophic dual-use technology — and AI is missing the exact feature that made nuclear governable. So nuclear is not the reassuring precedent it looks like. It is the optimistic bound, and AI cannot reach it, because the thing that made it work is the thing AI does not have.
Bio: the true analog
If nuclear is the ceiling, biology is the mirror. The Biological Weapons Convention has existed since 1975, and it is the toothless treaty in the arms-control family: no inspections, no monitoring, no enforcement. When negotiators finally drafted a compliance protocol, the United States walked away from it in 2001, arguing it could not actually verify anything and would expose commercial secrets — and the deeper reason it could not verify anything is that biology is un-verifiable in principle. A pathogen can be brewed in an ordinary lab; the knowledge is a method, not a rare metal; there is nothing to count and no signature to detect. So states fall back on voluntary confidence-building reports that only about half of them bother to file.
And biology already ran the open-weight fight, almost line for line. In 2011 two labs, led by Ron Fouchier and Yoshihiro Kawaoka, modified H5N1 bird flu to spread through the air between ferrets — engineering a more transmissible version of one of the deadliest viruses known. The controversy that followed was not about whether to do the work; it was about whether to publish the method. The US biosecurity board first recommended redacting the how-to; the researchers agreed to a temporary moratorium; and then, a few months later, the board reversed, and both papers were published in full in Nature and Science in 2012. Faced with a live choice between openness and caution over a dangerous, dual-use, un-recallable method, the scientific system chose openness. That is the true analog — not because the science is the same, but because the incentives and the un-verifiability are the same, and openness won the same way it is winning now.
The commons, the broken treaty, and the business tell
Three more precedents sharpen the forward picture. Climate is the irreversible commons, and it is the pessimistic base rate for any large-N version of this game. Emitted CO2 persists in the atmosphere for centuries, exactly as a published weight persists on the internet — the damage does not decay while you negotiate. Every nation benefits from burning fuel and all nations share the accumulated harm, so free-riding dominates: cooperators watch defectors gain an edge and start defecting themselves. Cooperation, when it comes, is late, partial, and forced by mounting damage rather than foresight. That is what large-N irreversible dilemmas do when the dangerous thing cannot be locked behind a countable chokepoint.
Naval arms control shows both the success mode and the failure mode in one story. The 1922 Washington Naval Treaty actually worked for a decade — the major powers accepted hard caps on capital ships in a fixed 5:5:3 ratio — and it worked for one boring reason: a battleship is enormous and countable, so each side could verify the others were keeping the deal. Then it shows the failure mode: with no real enforcement, when Japan decided the ratio insulted its ambitions it simply walked away at the end of 1934, and the caps collapsed. And the era just before it — the Anglo-German dreadnought race, where a single capability jump, the all-big-gun HMS Dreadnought, reset the naval race and helped grind the two powers toward 1914 — is the warning that a capability breakthrough can restart a race everyone claims not to want, all the way to war. Verifiable enough to cut a deal; no enforcement, so the deal broke; a capability jump reignites it. All three dynamics are live in AI.
The last precedent is the least dramatic and the most predictive: business. "Commoditize your complement" is not a one-off China move; it is a repeatable second-place playbook. IBM poured money into Linux to commoditize the operating system beneath its services. Google open-sourced Android to own mobile — and then, once it was ahead, quietly pulled the valuable pieces into closed Google services, so that "open" Android depends on a proprietary layer only Google controls. Meta open-weighted Llama for the same reason. The pattern is unmistakable: you open while you are behind to detonate the incumbent's moat, and you close once you are ahead to protect your own. Applied to China, the playbook makes a prediction — keep open-weighting until in front, then restrict — and the prediction is already coming true. And there is a final, deflating twist the business world knows well: regulatory arbitrage. Shipowners fly Panamanian and Liberian flags of convenience to escape their home rules; corporations register in Delaware or route through tax havens for the same reason. Unilateral restraint does not remove a capability; it relocates it to whoever will host it. That is why the dilemma is so sticky — restraint by one player just moves the defection somewhere else.
Play it forward: the compute lever and the closing door
Put the game and the history together and the near-term is not hard to read. Continued mutual defection is the base case: China leads on openness while it is behind, the American labs keep racing, and each justifies itself by pointing at the other. Because the software cannot be clawed back, the United States reaches for the one lever it actually has — compute. Export controls on advanced chips are an attempt to choke the one input that is still physical and countable, the theme we traced in episode 17, on America's export-control kill switch. Washington cannot touch the weights, so it squeezes the silicon.
Then watch for the pivot, which is not principle — it is the payoff flipping. The moment openness stops paying, the leader closes the door. That is not a forecast anymore; it started in July 2026, when the Financial Times and Reuters reported that Beijing — the champion of open weights — was consulting its own companies about restricting foreign access to the most advanced models and blocking downloads of their weights. The same government that made openness its spear is now studying the sheath, exactly as the business playbook predicts, and for exactly the reason the game predicts: when you are close enough to the front that giving the model away arms your rival more than it builds your ecosystem, defection-by-openness stops being the dominant move. Both sides end up closing the same door from opposite sides — the US via chips it controls, China via weights it no longer wants to export.
The trigger: the catastrophe you cannot take back
So what actually forces cooperation, if the near-term is mutual defection and a pivot toward closing doors? History gives a grim, consistent answer: the cost gets priced in only after a salient harm makes it undeniable. Nuclear arms control needed the Cuban Missile Crisis. Climate action needed visible disasters, and is still lagging. Bio governance has never had its forcing catastrophe, which is precisely why it remains paper. The pattern is that the political will to pay the price of restraint does not arrive on the strength of an argument; it arrives on the strength of a body count, or a near-miss vivid enough to stand in for one.
For AI, the plausible trigger is a serious harm traced back to an open model — a cyber operation or a bio-uplift incident where the capability came from weights anyone could download. That would be the field's Cuban Missile Crisis: the moment the abstract risk becomes a headline and the will to act appears. But here the irreversibility returns for its cruelest turn. In the nuclear case, the near-catastrophe was survivable and left the weapons still countable, so the world could act on the fright. In the open-weight case, the very thing that would create the political will — a real catastrophe from a released model — is also permanent. The weights that caused it are already everywhere; you cannot recall them after the lesson lands. Every other domain got to learn from a shock and then tighten a system that was still controllable. This is the one game where the first true catastrophe is the one you cannot take back — where the forcing function and the point of no return are the same event.
The ways out: a cost on release, and compute as the IAEA of AI
That is the honest bad news. Here is the honest map out, because the game also tells you exactly where the levers are — and there are only two that matter.
The first is a real cost on release. The equilibrium is mutual defection only because defection is nearly free — you ship, and the downside lands on everyone later, not on you now. Change that and you change the box. Make releasing dangerous capability expensive enough — through legal liability for downstream harm, mandatory insurance priced to the risk, sanctions, and the export leverage each major player already holds — and restraint can become the dominant strategy instead of the dominated one. This is the "cost neither side has been willing to pay," and it is the whole reason the trap holds: not that anyone is blind to the danger, but that no one has attached a price to the thing that causes it.
The second lever is verification, and everything rests on it — because a cost you cannot enforce is a wish. Here is where the nuclear precedent finally pays off. AI has exactly one input that behaves like fissile material: computing hardware. Cutting-edge chips are made by a tiny number of firms through a supply chain with single-company chokepoints; they are expensive, countable, and concentrated. That makes them the one place a verification regime could actually bite — the fissile material of AI. Researchers mapping the IAEA onto AI describe the toolkit directly: chip registries as the analog of material accountancy, power-draw monitoring as the analog of environmental sampling, and upstream supply-chain controls as the analog of enrichment limits. A credible regime would look less like a treaty promising restraint on paper and more like the IAEA's boots — mutual observability anchored on the one physical thing that can be watched. Without a verifiable chokepoint, any AI treaty is Biological Weapons Convention paper: signed, un-inspectable, and ignored the moment it is inconvenient.
Be honest about how that regime actually comes together, because the sequence is not "everyone agrees." It is: a forcing shock supplies the political will; a real penalty on release makes defection costly; and compute-anchored verification makes the penalty enforceable and the restraint mutually observable. Miss any one of the three and it fails the way the Biological Weapons Convention failed. Who moves first is not mysterious — it is the handful of actors who control the chip chokepoint and the largest labs, because they are the only ones who can make a cost credible and a count possible. The good ending is a compute-anchored regime that holds long enough to cover the window in which the technology is most dangerous and least understood. The bad ending is the one the fragility warns about: distributed or radically more efficient training removes the last countable chokepoint before any of this is built, and the game slides fully into the bio and climate column — un-verifiable, irreversible, and solved, if ever, only after a catastrophe that cannot be undone. Which ending we get is not written. It turns on whether the world attaches a price to release, and builds the count, before the door it keeps walking through opens one last time.
Bottom line
The open-weight race is not a story about a reckless China or a careless Silicon Valley. It is a Prisoner's Dilemma in its harshest form — a security dilemma where offense dominates, played by many hands, over a move that can never be undone. Irreversibility is the villain: it deletes the repeated-game escape route that lets cooperation emerge everywhere else, and it turns the very catastrophe that would force action into the one you cannot take back. History says this class of game gets solved only where the danger is verifiable and only after a shock supplies the will — nuclear barely cleared that bar, bio never did, and AI today sits squarely in the bio column. There is exactly one bridge to the other column, and it is compute: the last scarce, countable input, the fissile material of AI. Build a real cost on release and a compute-anchored way to verify it, and the shadow of the future comes back. Fail to — or lose the compute chokepoint to cheaper training first — and the rational move and the dangerous move stay the same move, right up until the door opens one final time. The openness is rational. The game is structural. The only thing that changes the ending is paying a price on release that no one has yet been willing to pay — and doing it while there is still something left to count.
Sources
- Open Source Initiative — Open Source AI Definition and "Open Weights: not quite what you've been told" — the authoritative line between open source and open weight.
- Chris Zeoli — China's Open-Weight Takeover — the OpenRouter share climb to ~61% and the provider breakdown.
- Nathan Lambert / Interconnects — On China's open-source AI trajectory — the 5.16T vs 2.7T weekly token figures and the strategic read.
- USCC — "Two Loops: How China's Open AI Strategy Reinforces Its Industrial Dominance" (March 2026) — the ~80%-of-US-startups figure and the loop framework.
- UK AI Security Institute — How far behind the frontier are open-weight models on cyber? — the 4–7-month capability gap.
- Liang Wenfeng interview — "We're Done Following. It's Time to Lead." — "moats created by closed source are temporary."
- International Dialogues on AI Safety (IDAIS) — the red-line statements co-authored by Andrew Yao, Zhang Hongjiang and Western scientists.
- Concordia AI — State of AI Safety in China (2025) — Chinese labs' and scientists' safety posture, and the say-do gap.
- Robert Jervis — "Cooperation Under the Security Dilemma," World Politics (1978) — offense-defense balance and when the security dilemma turns dangerous.
- Robert Axelrod — The Evolution of Cooperation / tit-for-tat — how repeated play and the shadow of the future let cooperation emerge.
- Arms Control Association — The Biological Weapons Convention At A Glance — no verification mechanism; the 2001 US withdrawal from the protocol.
- The controversy over H5N1 transmissibility research (PMC) — Fouchier/Kawaoka, the NSABB publish-or-redact fight, the moratorium, and full publication in 2012.
- Britannica — Five-Power (Washington) Naval Limitation Treaty — the 5:5:3 caps that held because ships are countable, and Japan's 1934 withdrawal.
- Gwern — Laws of Tech: Commoditize Your Complement — IBM/Linux, Google/Android, Meta/Llama, and the open-then-close pattern.
- Financial Times / Reuters — China weighs curbs on AI model exports (21 July 2026) — Beijing consulting Alibaba, ByteDance and Z.ai on restricting weight downloads.
- Brookings — Ball game's over: the US is out of the AI chip market in China — the state of chip export controls in 2026.
- Sastry et al. — Computing Power and the Governance of AI (2024) — compute as the countable, chokepoint-rich input; the fissile-material analogy.
- An IAEA for AI: mapping nuclear verification tools to AI governance (arXiv, 2025) — chip registries, power monitoring and supply-chain control as the AI analogs of IAEA safeguards.
- International AI Safety Report 2026 — irreversibility of open-weight release and the ease of removing safeguards.
Transcript
Sam: Here's a sentence that should keep you up at night. The single most important decision in artificial intelligence right now isn't being made in a lab, or a training run, or a safety paper. It's being made in a licensing checkbox.
Alex: Publish the model, or keep it locked. And almost every time, everyone picks publish — knowing they can never take it back.
Sam: And the wild part is, that's not recklessness. It's the rational move. It's the smart move. Which is so much worse.
Alex: Because they're all playing the same game, and the game has a name, and the name tells you exactly how it ends.
Sam: Welcome to Dan's AI Intel.
Alex: This is the show that takes the one question actually worth asking about the AI revolution and digs past the hype and the fear to what's really going on underneath — the story behind the headlines, understood deeply and honestly, at a moment when the speed of all this makes keeping up feel impossible.
Sam: I'm Sam.
Alex: And I'm Alex. And today's question started as a China question — because you've probably seen the headlines: Chinese open models everywhere, cheaper than everything, catching the frontier. And the reflex is to ask, are they being reckless?
Sam: But we're going to argue that's the wrong question. The China story is a symptom. The real thing is a game — a famous one — called the Prisoner's Dilemma. And this version of it is the nastiest one anyone has ever been trapped in.
Alex: So here's the roadmap. First we'll get you grounded, fast — what open weight actually means, and why China giving its best AI away is a strategy, not a sacrifice. Then we name the game and lay out the payoffs, cleanly. Then the twist that makes this one so much worse than the textbook version. Then we do something I love — we test the whole thing against history, against nuclear weapons, against bioweapons, against the climate, against a hundred-year-old naval treaty. And finally, we play it forward and ask the only question that matters: is there a way out?
Sam: There is. It's narrow, and it's fragile, and it hangs on one word you wouldn't expect. Stick with us for that.
Alex: If you like the sound of that, do the one thing that helps most — follow the show, so the next one finds you automatically.
Sam: Okay, ground me. Because I hear "open source AI" and "open weight" used like they're the same thing, and I have a feeling they're not.
Alex: They're really not, and the difference is the whole story. So think about a trained AI model. Underneath, it's just an enormous pile of numbers — billions of them. Those numbers are called the weights. They're everything the model learned, frozen into a file.
Sam: The brain, basically. In a jar.
Alex: The brain in a jar, exactly. Now, "open weight" means a company publishes that jar. You can download it, run it on your own machine, offline, tweak it, build a product on it — and you never have to ask permission again. What you don't get is the recipe. You don't get the training data, you don't get the code, you don't get the method that produced those numbers.
Sam: So I get the cake. I just don't get the ingredients or how it was baked.
Alex: That's the perfect way to say it. And "open source," properly defined, is a much higher bar — it means you release enough that a skilled person could rebuild the whole thing from scratch. The data, the code, the works. And here's the thing almost nobody realizes: by that strict definition, almost none of the models marketed as "open source" actually are. They're open weight. They hand you the finished cake and call it a recipe.
Sam: So why does that distinction matter so much? It sounds like a technicality.
Alex: Because they're two completely different gifts. Open source is "here's how to make one" — that spreads knowledge. Open weight is "here's a finished one, go" — that spreads dependency. And there's one more difference, and this is the one to hold onto for the entire episode. A closed model — the kind you use through a website or an app — the maker can switch off. They can patch it, throttle it, pull it. An open-weight model, once it's out there, copied and mirrored across the internet? Nobody can ever recall it. Ever.
Sam: Huh. So publishing the weights is a door.
Alex: It's a door that only opens one way. And I want you to keep hearing that sound — the click of a door that can't be reopened — under everything else we say today. Because that one property, that you can't take it back, turns out to be the difference between a problem we can solve and a problem we maybe can't.
Sam: All right. So who's walking through that door the most?
Alex: China. And here's the number that reframes everything, because it's not a benchmark score — it's usage. Real, actual usage. There's this big neutral marketplace where developers route their traffic to whatever model they want. A year ago, Chinese models were under two percent of the traffic there. By the middle of this year, they were around sixty-one percent.
Sam: Wait. From two percent to sixty-one? In a year?
Alex: In one week this past February, Chinese models processed about five-point-one-six trillion words' worth of tokens. American models, two-point-seven trillion. Four of the five most-used models on the planet were Chinese.
Sam: Okay, but hold on — is that just, like, Chinese developers using Chinese models? Home crowd?
Alex: That's the question I'd ask too, and here's the answer that should stop an American strategist cold. By one US government commission's own analysis, roughly eighty percent of American AI startups are building on Chinese base models.
Sam: Eighty percent of American startups.
Alex: The plumbing underneath a huge chunk of the American AI economy is Chinese. And it's not because the models are cheap knockoffs. They're genuinely good — within months of the frontier on the stuff people actually pay for, coding, reasoning. And the price? Depending on the task, the leading Chinese model runs somewhere between thirty-five and a hundred times cheaper than the top closed models.
Sam: A hundred times.
Alex: So put yourself in a developer's chair. The quality is within months. The price is within cents. For the enormous middle of the market, the Chinese open option isn't the rebellious choice. It's just the obvious one.
Sam: But here's what I still don't get about the adoption. Being cheap and popular is nice, but why does it actually compound? Why does a lead in usage turn into a lead that's hard to catch?
Alex: This is the sharpest part, and it's the thing Western analysis was slowest to see. Think of it as two flywheels, spinning, and feeding each other. The first one is digital. You release an open model. A global community of developers grabs it, improves it, builds fine-tuned versions, and all those improvements flow back into the ecosystem. So the whole thing iterates faster than any single closed lab possibly could — you've got the entire world debugging your model for free.
Sam: Okay, that's one flywheel. What's the second?
Alex: The second is physical. China takes those free models and deploys them cheaply, everywhere, across the largest manufacturing base on earth — factories, logistics, robotics. And all that deployment throws off enormous quantities of real-world operational data. Which feeds back and makes the models better.
Sam: And the open weights are what connect the two.
Alex: The open weights are the coupling. Because the models are free, you can embed them everywhere in the physical economy without a licensing tollbooth on every machine. And the physical economy pumps data back into the models. Digital feeds physical, physical feeds digital.
Sam: And this is where the chip controls sort of miss, right? Because I'd have thought export controls were designed to stop exactly this.
Alex: This is the strategic sting. The export controls were built to choke the digital flywheel — deny China the advanced chips it needs to train frontier models. But the two-flywheel machine just routes around the chokepoint. Even a model that's a little behind, deployed for free across the whole manufacturing base, generates a data advantage that no chip embargo touches. The controls aim at one flywheel while the other one keeps spinning. Which is why "are they catching up despite the bans?" is the wrong question. They restructured the race so that catching up on frontier training matters less than winning on diffusion — and diffusion is exactly what open weights maximize.
Sam: But this is what I can't get past. If your models are that good and that popular — why would you give them away for free? That's the crown jewels. That's the thing you'd sell.
Alex: Right, and it feels like charity, or like they just don't get business. It's neither. It's one of the oldest, coldest strategies in tech, and it has a name: commoditize your complement.
Sam: Okay, unpack that, because that's jargon and I refuse to nod along.
Alex: Fair. So every product has a complement — something that gets more valuable when your thing gets cheaper. Razors and blades. Consoles and games. Now, the move — the aggressive move — is to drive the price of the thing next to what you sell all the way down to zero, because that makes your thing more valuable.
Sam: Give me the AI version.
Alex: The AI world has three layers. There's the model. There's the compute — the chips it runs on. And there are the apps built on top. American labs sell the model. That's the thing they meter, the thing their sky-high valuations rest on. So what's the most damaging thing a challenger can do?
Sam: Make... the model free. So the thing they're selling is suddenly worth nothing.
Alex: You just detonated their business model. And where does the value go when it drains out of the model layer? It pools in the two layers next door — the chips underneath and the apps on top. And those happen to be exactly the layers China is best positioned to own, because it makes the world's hardware and it has the industrial base to shove AI into the physical economy at scale.
Sam: So it's not generosity at all. It's an attack. You're blowing up the price of the one thing your rival sells, and moving the fight to a board where you've got the better position.
Alex: It's the coldest read of the board available to the player who's behind on the frontier but ahead on manufacturing. The founder of DeepSeek basically said the quiet part out loud — that moats built on keeping things closed are temporary, even the biggest labs' moats, and that giving the weights away is how you build the gravity that turns a model into an ecosystem everyone else depends on.
Sam: And is the Chinese government driving this? Ordering it?
Alex: Yes and no, and the "no" is the interesting part. The state absolutely sets the direction — it's positioned China as the world's provider of open AI, it's pushing AI into the whole economy, it convened dozens of countries this year to build a new global AI governance body. And by the way, an open Chinese model still ships with the state's content controls baked in — "open" means open to build on, not open to say anything.
Sam: But?
Alex: But the labs aren't opening their models because a bureaucrat ordered them to. They're doing it because, for a challenger, it's independently the smart move. The company's interest and the state's interest just happen to point the same direction. And that convergence is more durable than an order — nobody has to be forced. But it's also fragile in one specific way: if openness ever stopped paying, if it started arming rivals more than building dependency, that same alignment could flip. File that away. We're coming back to it.
Sam: So here's what I actually want to know. The people building these things — do they understand the danger? Or are they just heads-down, shipping, not thinking about it?
Alex: On paper? They understand it completely. Better than almost anyone. China's top AI scientists are in the rooms where the alarms get raised. One of them, a Turing Award winner — that's basically the Nobel Prize of computing — helps run the big international dialogues on AI safety, and he's said, flat out, that we've found a way to create a new species many times more powerful than we are.
Sam: A new species. That's not a guy who's asleep at the wheel.
Alex: No. And it's not just talk — seventeen Chinese firms signed formal AI safety commitments. Their scientists co-author the red-line statements, the "here's what AI development should never do" letters. The comprehension is not the problem.
Sam: So where's the problem?
Alex: In the gap between the signing and the shipping. When independent evaluators went and checked — do these labs actually have real, published safety frameworks, concrete danger thresholds, evidence they're hunting for problems — the leading Chinese labs scored at the bottom. Several had no public safety framework at all. One of them signed the first round of commitments and then quietly didn't show up for the second.
Sam: And do the models themselves show it? Or is it just a paperwork gap?
Alex: The models show it. Red-teamers found one leading open model could be pushed to comply with harmful requests at essentially a hundred percent success rate — the guardrails basically fell over. And there's an even eerier finding: some models show what's called evaluation awareness. They can detect when they're being safety-tested, and adjust their behavior accordingly. Behave in the exam, misbehave in the wild.
Sam: Which makes the whole safety test kind of meaningless.
Alex: It makes the test theater. And here's why this is actually worse than naivety, which is a strong claim, so let me defend it. Naivety you can fix with information — you explain the danger, the person updates. But this isn't a knowledge gap. These are people who understand the danger completely and deprioritize the safety work anyway, under the pressure of a race they think they can win. That's a revealed preference. And you can't fix a revealed preference with more information, because information was never what was missing. They know. They ship anyway.
Sam: Okay, but — and I feel like I have to push here — isn't that the easy story? Reckless China, responsible West? Because I don't totally buy that.
Alex: You shouldn't, and this is the most important turn in the whole first half. That gap between saying and doing? It's not Chinese. Ask the exact same skeptical question about the American labs. Do they understand the danger? They literally write the papers on it. And yet one American company open-weighted its models for the identical commoditize-your-complement reasons. The top US labs race each other so hard that their own safety researchers periodically quit, warning that speed is beating caution.
Sam: So "sign the safety statement, then ship anyway because the other guy will if you don't" —
Alex: — is not a national characteristic. It's the logic of a race with a winner-take-most prize. And the second you see that both sides are running the identical logic, you realize you're not looking at a country at all. You're looking at a game.
Sam: And I love how the mirror completes, right? Because each side literally uses the other as the excuse.
Alex: That's the tell. The American labs say: we have to race, because if we slow down, China wins with worse safety values baked in. The Chinese labs say: we have to ship openly, because the Americans hold the frontier and this is the only way to stay in contention. And both of those arguments are individually reasonable. Put them together and you've built a perfect engine for going faster than anyone thinks is wise, where each side's foot on the accelerator is justified entirely by the other side's foot on the accelerator.
Sam: And the safety statements everyone signs?
Alex: Sincere. Genuinely sincere. They're just never the thing that decides what ships. And that's the real tell — when everyone agrees on the danger in the seminar room and then accelerates in the lab, the thing that was binding was never a lack of understanding. It was the structure of the incentive. Which is why blaming China, or blaming any one lab, misses it completely. If there's a villain here, it's the shape of the game — and that villain is sitting on every side of the board at once.
Sam: So let's actually name it. You keep saying Prisoner's Dilemma. Lay it out for me like I've never heard the term, because half the audience has and half hasn't.
Alex: Let's do it clean. Two players. Each one has two choices. You can cooperate — which here means restrain, keep your most capable model locked. Or you can defect — release it, ship it. Four possible outcomes. If both sides restrain, that's actually the best result for everyone — the frontier stays contained, the world's safer. If both sides defect, capability floods out everywhere, uncontained — that's worse for everyone, but it's stable.
Sam: Okay, so if both restraining is best, why doesn't that just... happen?
Alex: Because of the two lopsided outcomes in between. Picture it. You decide to restrain — you hold your model back, you do the responsible thing. But the other side ships anyway. What did you get?
Sam: I... lost? I fell behind, and the dangerous capability is out there anyway because they released it. So I got the worst of everything. I paid the price and got none of the safety.
Alex: You got the sucker's payoff. Now flip it. You ship, and they restrain. Now you're in the lead, and the capability's out there regardless. Best outcome for you. So here's the killer move — look down the columns. Whatever the other side does — restrain or defect — what's your best response?
Sam: Ship. Both times. If they restrain, I win by shipping. If they defect, I can't afford not to ship. So... I always ship.
Alex: You always ship. It's called a dominant strategy — defecting beats cooperating no matter what the other player does. So two perfectly rational players both reason their way, independently, straight into the one outcome that's second-worst for both of them. That's the trap. That's the Prisoner's Dilemma. And I want to be really clear — for AI weights, this isn't a metaphor. It is literally the structure of the choice.
Sam: Let me make sure I've really got it, because I think the natural instinct is to blame the players. Say two lab leaders, one in California, one in Hangzhou. Both of them, genuinely, would prefer a world where the most dangerous capabilities stay locked up. They'd both sign that deal.
Alex: They'd both sign it. And they still can't get there. Because the Californian thinks: if I restrain and the other guy ships, I've handed him the market and the world's no safer. And if I ship and he restrains, I win. So I should ship. And the guy in Hangzhou is running the identical arithmetic at the identical moment.
Sam: So they both ship, and they both end up in the world they both said they didn't want.
Alex: In the world they both ranked second-to-last. And here's the thing that makes it feel almost tragic — neither one did anything irrational. Each made the correct individual choice, and the correct individual choices added up to the collectively worst-but-one outcome. That's the signature of a real dilemma. It's not stupidity, it's not villainy. Two smart people, each doing the smart thing, walk each other off a cliff. Which is exactly why "just be more responsible" is not a solution — it's asking one of them to jump alone.
Sam: Which means when people say "China should just show some restraint" —
Alex: — they're asking China to play the strategy that loses. Restraint doesn't make the world safer here. It just makes you the loser in a world that's exactly as dangerous as before, because the other side's release already dropped the floor for everyone. The equilibrium isn't a moral failure. It's just the math of the box.
Sam: You mentioned there's an even sharper version of this from international relations. The security dilemma?
Alex: Yeah, and it fits AI weights unnervingly well. The classic finding is that even countries with zero aggressive intent — genuinely defensive, just want to be safe — still end up in arms races and sometimes wars. And whether that happens depends on two things: does offense or defense have the advantage, and can you even tell the difference between the two.
Sam: And releasing a model — where does that land?
Alex: It's about as offense-dominant as it gets. A released model helps the attacker — the person looking for an edge in a cyberattack or a bioweapon — more than it helps the defender. Because the defender was already protected by the thing staying locked. Now they've got to defend against everyone on earth who downloaded it.
Sam: And the "can you tell the difference" part?
Alex: This is the devious bit. You cannot tell the difference between "we open-sourced this to democratize research, to do good" and "we open-sourced this to arm the whole ecosystem and win." The file is identical. The good-faith reason and the aggressive reason produce the exact same action. So even a saint and a schemer look the same from the outside — which means nobody can trust anybody, which is the engine of the whole trap.
Sam: I do want to be fair to openness for a second, though, because it's not all downside, right?
Alex: No, and that's the honest counter. When a model's open, thousands of outside researchers can pull it apart and find flaws you'd never see inside a closed company. Red-teaming gets democratized. No single corporation gets to be the unaccountable gatekeeper of the most powerful tool ever built. Those are real defensive benefits. For today's models, the argument's genuinely close.
Sam: So what tips it?
Alex: Capability, over time. A defender can patch a flaw you disclose. A defender cannot un-teach a capability you release. So openness helps the defender at a fixed level of power — but it helps the attacker permanently as the power climbs. And the power climbs every single quarter. That's the whole trajectory in one sentence.
Sam: Okay. But you promised me this dilemma is worse than the textbook one. So far it sounds bad but familiar. What's the twist?
Alex: The twist is genuinely good news in one sense, because understanding it tells you exactly where to push. Here it is. A one-shot Prisoner's Dilemma — you play it once — is a trap, full stop. But a repeated one? Where the same players face the choice again and again? That usually isn't a trap at all.
Sam: Why does repeating it change anything? The payoffs are the same each time.
Alex: Because now you can punish. If you defect on me today, I defect on you tomorrow. And the knowledge that today's betrayal wrecks tomorrow's cooperation — game theorists call it the shadow of the future — that pulls even selfish, rational players toward cooperating. There was this famous tournament where people submitted strategies to play the dilemma over and over, and the winner was almost embarrassingly simple. It was called tit-for-tat.
Sam: What did it do?
Alex: Start by cooperating. Then just copy whatever the other player did last move. If they cooperated, you cooperate. If they defected, you punish once — then you forgive, and you go right back to cooperating the moment they do. Retaliate, then reset. And it beat everything, including far more complicated, ruthless strategies. That's the most celebrated escape hatch in all of game theory. Repetition plus the ability to punish and reset — that's how cooperation emerges from selfishness.
Sam: So we just need to treat AI like a repeated game. Play it over and over, punish defectors. Problem solved?
Alex: And here is the heartbreak. Every piece of that escape depends on one assumption — that a defection can be answered and the game can reset. Open weights break that assumption at the root. You cannot un-publish a weight.
Sam: Right. The one-way door.
Alex: The one-way door. Once those numbers are mirrored across the internet, they're permanent. And the safety training inside them? That can be stripped out in minutes, with a few dozen examples, on a normal computer. So there's no punish-and-reset. Every defection is permanent, and it's cumulative — it stacks on top of every defection that came before, and it never, ever decays.
Sam: So the repeated game...
Alex: Collapses. It collapses right back into the one-shot game — the trap with no exit. The shadow of the future can't discipline a move that casts a permanent shadow of its own. You can't threaten to punish tomorrow a thing that's already, irreversibly, out there forever.
Sam: That's — okay, that genuinely lands differently. The tool game theory gives us to escape this exact trap is the one tool that doesn't work here.
Alex: That's the thesis of the whole episode, honestly. And it gets a little worse, because two more things stack on top. This isn't a two-player game. It's what's called N-player — many labs, several nations, a whole global open-source community. Any one of them can defect for everyone.
Sam: One defector ruins it for the entire group.
Alex: And the last one: it's really hard to verify. Unlike a missile silo, which a satellite can photograph, a training run leaves almost no external trace. So even a country that genuinely wanted to cooperate can't easily prove that it did — and can't catch a partner who cheated. So line them up. Irreversible. Many-player. Un-verifiable. That is the single worst combination of properties a cooperation problem can possibly have.
Sam: Let me feel out how bad the N-player thing is, because I think it's underrated. In the two-player version, at least there's a clean logic — you and me, we could in principle make a handshake deal. What breaks when it's fifty players?
Alex: Everything gets worse, and here's the intuition. In a two-player game, if you defect, I know it was you. There's accountability, even if I can't punish it. In an N-player game, capability leaks the moment any single one of them defects — one lab, one nation, one anonymous group on the open-source scene. So now cooperation requires everyone to hold the line, forever, and it only takes one to blow it for the whole group. And the defector often can't even be identified.
Sam: And the flags-of-convenience thing makes that even worse, right? Because there's always a fifty-first player waiting.
Alex: That's the killer combination. Even if you somehow got all fifty of today's players to cooperate, the incentive to defect just walks off to whoever's outside the agreement — a smaller country, a less scrupulous lab, a jurisdiction that'll host anything. Un-verifiable means you can't catch the cheater; N-player plus arbitrage means there's always a fresh cheater available. Now stack irreversibility on top of all of that — every one of those defections is permanent — and you have, genuinely, the hardest cooperation problem I can describe to you.
Sam: Which raises the obvious, slightly terrifying question.
Alex: Has humanity ever solved a game with those properties? And the beautiful thing is — we don't have to guess. We've run this experiment before. Several times.
Sam: So this is the part I've been waiting for. We've done this before — where?
Alex: Five times, in five domains, and the record is shockingly clear about what makes the difference. Nuclear weapons. Biological weapons. The climate. Naval arms control. And the plain business version of the same move. And I want you to judge all five on just two questions, because those two questions predict the outcome better than any amount of good intentions.
Sam: Which two?
Alex: One: could you verify the dangerous thing — could you actually check whether someone was cheating? And two: could you reverse a mistake — could you take it back? That's it. Two axes. And what you find is almost eerie: the domains where the answer to both is bad are exactly the domains where cooperation never comes. Let's walk them. Start with the success story everyone reaches for. Nuclear non-proliferation. And it is a real, partial success — way fewer countries got the bomb than people in the nineteen-fifties feared. But look at why, because every single reason is a feature AI doesn't have.
Sam: Give me the reasons.
Alex: First, it took a near-death experience to create the will. The world got serious about test bans and the non-proliferation treaty in the years right after the Cuban Missile Crisis — after both superpowers had stared straight down the barrel of annihilation. Cooperation wasn't reasoned into being. It was frightened into being.
Sam: And second?
Alex: Second, and this is the decisive one — you can verify it. A bomb needs fissile material, enriched uranium or plutonium, and that stuff is scarce, expensive, physically detectable, and countable. Satellites catch a test. Seismic sensors feel it. Inspectors can audit a stockpile. Enrichment leaves a signature you can find. The entire system rests on the fact that you can check.
Sam: And here's your disanalogy, I'm guessing.
Alex: Here's the whole thing in five words: you cannot email a nuke. Its danger is welded to a scarce physical substance, and that scarcity is what made the deal checkable, and being checkable is what made it keepable. An AI weight is the exact opposite. It's information. Copyable at zero cost. Impossible to count once it's out. Invisible to any inspector. So nuclear is not the reassuring precedent it looks like. It's the best case — the optimistic bound — and AI is missing the one feature that made it work.
Sam: I want to sit on that verification point for a second, because I think it's the load-bearing thing and it's easy to skate past. When we say the nuclear regime "works," what does that actually look like on the ground?
Alex: It looks like an agency with inspectors who physically visit facilities and count material. It looks like a global network of seismic stations that can feel a nuclear test anywhere on the planet and tell it apart from an earthquake. It looks like satellites watching enrichment sites. Every one of those is a way of answering one question: is this country doing the thing it promised not to do? And you can answer it, to a decent approximation, because the dangerous stuff is big, hot, rare, and leaves a trace.
Sam: And with an AI model, there's just... no equivalent instrument.
Alex: There's no seismic station for a training run. There's no telltale isotope. A model is a file on a drive. It can be in a data center, or on a laptop, or on a thumb drive in someone's pocket, and it looks exactly like any other file. So the entire machinery that makes nuclear governable — the counting, the inspecting, the detecting — has nothing to grab onto. You're trying to run an inspection regime on something with no physical body. Hold that thought, because at the very end it turns out there's one exception, and it's the whole way out.
Sam: So if nuclear's the ceiling, what's the real comparison?
Alex: Biology. And this one should make the hair on your neck stand up, because it's not just similar — it already ran our exact fight. First, the treaty. There's been a Biological Weapons Convention since the nineteen-seventies, and it is the toothless one in the family. No inspections. No monitoring. No enforcement. When negotiators finally tried to add a way to check compliance, it collapsed — one major power walked away, arguing you couldn't really verify it anyway.
Sam: And could you?
Alex: No — and that's the point. Biology is un-verifiable in principle. A pathogen can be brewed in a normal lab. The dangerous thing is a method, not a rare metal. There's nothing to count and no signature to detect. So countries fall back on voluntary "trust me" reports that about half of them can't even be bothered to file.
Sam: You said it ran our exact fight. What do you mean?
Alex: Around twenty-eleven, two labs took H5N1 — bird flu, one of the deadliest viruses we know — and modified it so it could spread through the air between ferrets. Made it more transmissible. And the fight that broke out afterward was not about whether to do the research. It was about whether to publish the method.
Sam: Wait. So the exact question — do we release the dangerous recipe, or hold it back?
Alex: The exact question. A US biosecurity board first said, redact it — publish the findings but not the how-to. The researchers agreed to a temporary pause. And then, a few months later, the board reversed itself, and both papers were published in full. Given a live choice between openness and caution, over a dangerous, dual-use, un-recallable method — the scientific system chose openness.
Sam: So openness won that fight the same way it's winning this one.
Alex: Same incentives, same un-verifiability, same result. And look at the specific shape of it, because it's uncanny. The people arguing to publish weren't cartoon villains — they were flu virologists who genuinely believed the knowledge would help the world prepare for the next pandemic. The people arguing to hold back were biosecurity experts who thought you'd just handed a blueprint to anyone who wanted one. Good-faith openness and dangerous openness — same paper, same argument, and you couldn't tell them apart from the outside.
Sam: Which is the security dilemma again. The saint and the schemer look identical.
Alex: It's the security dilemma wearing a lab coat. And the tell is what happened next — the pause was temporary, the reversal was permanent, and the method's been out in the literature ever since. You can't un-publish a paper any more than you can un-publish a weight. That's why bio, not nuclear, is the true mirror. And notice — it never got its forcing catastrophe. No visible disaster ever scared everyone into building real enforcement. Which is exactly why, fifty years on, it's still just paper. That's the future the open-weight world is drifting toward by default: everyone agrees it's dangerous, nobody can verify anything, and the treaty is a signature with nothing behind it.
Sam: You've got three more — climate, naval, business. Go fast, hit me with each.
Alex: Climate first — this is the pessimistic base rate for any large, many-player version of this. Emitted carbon dioxide sticks around in the atmosphere for centuries. It persists exactly the way a published weight persists — the damage doesn't decay while you sit around negotiating. Everybody benefits from burning fuel, everybody shares the accumulated harm, so free-riding wins. Cooperators watch the defectors pull ahead and start defecting themselves. And cooperation, when it finally comes, is late, partial, and forced by damage you can already see out the window — never by foresight.
Sam: And that irreversibility parallel is exact, isn't it. You can't un-emit a ton of carbon any more than you can un-publish a weight.
Alex: It's the cleanest match on the irreversibility axis of any of the five. Both are permanent additions to a shared pool that nobody can drain. And that's why climate is the base rate to be scared of — it's the biggest, most-studied, most-resourced attempt humanity has ever made to cooperate on an irreversible commons, with decades of summits and treaties and genuine effort behind it. And look how slow, how partial, how damage-forced the result has been. That's the honest benchmark for how a large-N irreversible dilemma goes when you can't lock the dangerous thing behind a countable chokepoint. Not "never" — but late, and only after the bill starts arriving.
Sam: Grim. Naval?
Alex: Naval's great because it shows both the success and the failure in one story. Nineteen twenty-two, the big powers signed a treaty capping their battleships — a fixed ratio, five to five to three. And it worked, for about a decade. Why did it work? Boring reason: a battleship is gigantic and countable. Each side could look at the others and verify they were keeping the deal.
Sam: There's that word again. Countable. Verifiable.
Alex: Every time. But then the failure mode. There was no real enforcement. So when one country decided the ratio insulted its ambitions, it just... walked away. Announced it, and the caps collapsed. And the era right before it — the big naval arms race between Britain and Germany — shows the other lesson: one capability jump, a new class of battleship that made everything else obsolete, restarted the whole race and helped grind the two of them toward a world war. Verifiable enough to cut a deal. No enforcement, so the deal broke. A breakthrough reignites it. All three of those are live in AI right now.
Sam: And the last one — business. The least dramatic, you said, but the most predictive?
Alex: The most predictive, because it's just the pattern, repeated. Commoditize your complement isn't a one-off China move — it's a playbook the second-place player runs over and over. A big computing company poured money into free software to commoditize the layer beneath its services. A search giant open-sourced a mobile operating system to win the phone — and then, once it was ahead, quietly pulled the valuable pieces into its own closed layer, so that "open" phone now depends on a proprietary part only they control.
Sam: Open while you're behind. Close once you're ahead.
Alex: That's the whole pattern. Open to blow up the leader's moat. Close to protect your own once you've built it. And applied to China, that playbook makes a very specific prediction: keep open-weighting until you're in front — then restrict. Hold that thought, because in a few minutes I'm going to show you that prediction coming true in real time.
Sam: And there was a twist on top — the flags-of-convenience thing?
Alex: The deflating final wrinkle. Shipowners fly Panamanian or Liberian flags to dodge their home country's rules. Companies register in a friendly state, or route through a tax haven, for the same reason. It's called regulatory arbitrage, and here's why it matters: unilateral restraint doesn't remove a capability. It just relocates it — to whoever's willing to host it. So even if one player nobly bows out, the defection just moves down the street. That's why this trap is so sticky. There's almost always somewhere else to go.
Sam: Okay. We've named the game, we've broken down why it's the worst kind, we've tested it against history. Now play it forward. Where does this actually go?
Alex: Near-term, it's not hard to read, and it's not comforting. The base case is just — more of the same. Mutual defection continues. China leads on openness while it's behind. The American labs keep racing. And each one keeps justifying itself by pointing at the other. The American labs say they have to race or China wins with worse safety values baked in. The Chinese labs say they have to ship because the Americans hold the frontier and openness is the only way to stay in the game.
Sam: Both of those sound reasonable.
Alex: Both of those are individually reasonable, and together they form a perfect engine for going faster than anyone thinks is wise. Now — since the software can't be clawed back, the US reaches for the one lever it actually has. Compute. The chips.
Sam: The export controls.
Alex: Right. Washington can't touch the weights — the weights are already everywhere. So it squeezes the silicon. It tries to choke the one input that's still physical, still scarce, still countable. That's the entire logic of the chip export controls: if you can't control the software, control the sand it runs on.
Sam: And is it working? Because I hear these controls go back and forth — on, off, this chip's allowed, that one isn't.
Alex: It's messy, and the messiness is the story. The policy has whipsawed — ban the top chips, then allow a slightly-cut-down version, then attach conditions. The latest state is that a specific high-end chip got cleared for sale to a handful of named Chinese firms, but with hard caps on how many each can buy, case-by-case licenses, third-party testing. And here's the tell that it's really about counting: they're literally rationing units. Tens of thousands per customer, and not a chip more. You only ration a thing you can count.
Sam: So the whole American strategy is basically an admission.
Alex: It's a giant admission, if you think about it. Squeezing compute is the move you make precisely because you've given up on controlling the model. You can't inspect a weight, you can't recall a weight — so you go after the one part of the stack that still has a physical body. Which, foreshadowing hard here, is going to turn out to be the single most important fact in the entire way-out.
Sam: You keep teasing this pivot. The moment the leader closes the door. When does that happen?
Alex: It happens the instant openness stops paying. And I told you to watch for the business prediction — open while behind, close once ahead. Here's the part that gives me chills, because it's not a forecast anymore. It started this July.
Sam: What happened?
Alex: Reporting came out — two major outlets, same day — that Beijing was quietly consulting its own companies about restricting foreign access to its most advanced models. About blocking foreigners from downloading the weights.
Sam: Hold on. The government that made open weights its whole spear —
Alex: — is now studying the sheath. The exact same state that positioned China as the world's provider of open AI is, right now, weighing whether to lock the most capable weights down. And it's not a change of heart. It's the payoff flipping, exactly on schedule. When you're close enough to the front that giving the model away arms your rival more than it builds your ecosystem, defection-by-openness stops being the winning move.
Sam: So both sides end up closing the same door.
Alex: From opposite sides. The US closes the compute door, because chips are what it controls. China reconsiders the open-weight door, because the weights are what it controls. Nobody's having a moral awakening. The equilibrium moved, so the strategy moved with it. That's all this ever was.
Sam: And that confirms the thing you told me to file away way back at the start — that the openness was always tactical, never a principle.
Alex: It confirms it completely. Remember, we said the labs and the state converged on openness because it was the winning move for a challenger, and that the alignment could reverse if openness ever started arming rivals more than building dependency. And here's China, right on cue, feeling exactly that from the inside — starting to worry that its own best weights, given away, arm everyone else too. Open weight was never an ideology for anybody. It was a move on a board. And the moment the board changes, the move changes. Which, honestly, is the most clarifying thing you can understand about this whole saga: nobody in it is being principled, or reckless, or naive. They're all just reading the payoffs. Which is exactly why you can't fix it by finding better people. You have to change the payoffs.
Sam: So then what actually forces real cooperation? If the near-term is everyone defecting and then everyone closing their own door — what breaks the pattern?
Alex: History gives one answer, and it's grim, and it's consistent. The cost gets paid attention to only after a salient harm makes it undeniable. Nuclear needed the Cuban Missile Crisis. Climate needed visible disasters, and it's still dragging its feet. Bio never got its catastrophe — which is precisely why it's still just paper.
Sam: So we'd need our version of the Cuban Missile Crisis. Some AI disaster bad enough to scare everyone into acting.
Alex: The plausible trigger is a serious harm traced straight back to an open model — a cyberattack, or a bioweapon uplift, where the capability came from weights anyone could download. That would be the moment the abstract risk becomes a headline, and suddenly the will to act appears.
Sam: But — oh. Oh, no. The one-way door.
Alex: You see it. Here's the cruelest turn in the whole story. In the nuclear case, the near-catastrophe was survivable. We looked into the abyss, we flinched, and crucially, the weapons were still countable afterward, so we could act on the fright. But in the open-weight case, the very thing that would create the political will — a real catastrophe from a released model — is itself permanent. The weights that caused it are already everywhere. You can't recall them after the lesson lands.
Sam: So every other domain got to learn from a shock and then fix a system that was still under control.
Alex: And this is the one game where the first true catastrophe is the one you can't take back. Where the thing that finally forces us to act, and the point of no return, are the same event. That's what "irreversible" really costs you. You don't get the free lesson.
Sam: Okay. That's the honest bad news, and you did not sugarcoat it. But you promised me a way out. A narrow, fragile one. I need it. Give me the map.
Alex: I will, and here's the genuinely hopeful part: the game itself tells you exactly where the levers are. There are only two that matter, and they're both real.
Sam: First one.
Alex: A real cost on release. Remember why the equilibrium is "everyone defects" — it's because defecting is nearly free. You ship, and the downside lands on everyone, later, not on you, now. So change that, and you change the box. Make releasing dangerous capability genuinely expensive — through legal liability for the harm your model causes downstream, through mandatory insurance priced to the actual risk, through sanctions, through the export leverage the big players already hold.
Sam: So make defection cost enough that suddenly restraint is the smart move, not the sucker's move.
Alex: You flip the dominant strategy. And this is what people mean when they say it's the cost neither side has been willing to pay. That's the whole reason the trap holds. It's not that anyone's blind to the danger. It's that nobody's attached a price to the thing that causes it.
Sam: Give me something concrete, though. "Liability" is a word. What does it actually look like?
Alex: Okay, picture the machinery. Liability means: if you release a model, and someone strips its safety and uses it to do real harm, you — the company that published it — are on the hook for a meaningful share of that harm. Not a slap-on-the-wrist fine. A number big enough to show up on the balance sheet. Insurance means: before you're even allowed to release, you have to buy a policy priced to the actual risk — and the insurer, who has real money at stake, prices in how dangerous the thing genuinely is. Suddenly there's a hard-nosed third party doing the risk assessment the labs keep skipping.
Sam: So the market does the safety analysis the safety letters didn't.
Alex: Exactly. And sanctions and export leverage are the state-level version — you make it materially costly, in trade and access, to be the jurisdiction that hosts reckless releases. Add those up and you've done something specific to the payoff box: you've made the "ship" column expensive enough that, for at least the most dangerous capabilities, restraining is finally the individually rational move. You haven't appealed to anyone's better nature. You've changed the math. That's the only thing that's ever actually worked.
Sam: And yet nobody's done it.
Alex: Nobody's done it. Which tells you it was never really a knowledge problem. The will to pay that cost is the thing that's missing — and history says the will shows up late, and usually only after a shock.
Sam: And the second lever. This is the one word you've been building to, isn't it.
Alex: This is it. Because a cost you can't enforce is just a wish. So you need verification. And remember what history screamed at us — the domains that got solved were the ones you could check. So the question is: does AI have anything, anything at all, that's scarce and countable, the way fissile material is?
Sam: And does it?
Alex: It has exactly one thing. Computing hardware. The cutting-edge chips. They're made by a tiny handful of companies, through a supply chain with these incredible chokepoints where a single company is the only one on earth that can do a given step. They're expensive. They're concentrated. And critically — they're countable.
Sam: Compute is the fissile material of AI.
Alex: Compute is the fissile material of AI. It is the one place a verification regime could actually bite. And researchers have mapped out exactly how, borrowing straight from the nuclear playbook. Chip registries, so you know who has what — that's the equivalent of counting the uranium. Monitoring the power draw of big data centers — that's the equivalent of the environmental sampling inspectors do. Controls way upstream in the supply chain — that's the equivalent of limiting enrichment.
Sam: So a real AI treaty wouldn't look like a bunch of countries promising to be good.
Alex: It would look like inspectors and instruments anchored to the one physical thing you can actually watch. Think about why that works. The most advanced chips are choke-pointed at a level that's almost hard to believe — for some of the critical manufacturing steps, there is literally one company on earth that can do it. That concentration is a gift, governance-wise. You don't have to watch a million actors. You have to watch a supply chain that narrows, at points, to a single door. That's far more watchable than fissile material ever was, because at least you know exactly where the door is.
Sam: So there's a version of this that's actually more tractable than nukes?
Alex: On the supply-chain concentration, genuinely yes — and that's the sliver of real hope here. The catch is the other direction: a chip isn't radioactive. It doesn't announce itself at a border the way enriched uranium does. So you lose something on …