Dario Amodei’s argument, stripped to the claims. Not a pause. Not a race to the bottom. Slow the rate at which models get more capable so safety work can keep up — while the U.S. still holds the lead.
Amodei has spent twelve years arguing AI can raise the quality of human life — disease, growth, abundance, a “renaissance of democracy.” He still believes that. What changed this summer is the slope.
Pacing means companies take enough time to align and safeguard models, and third parties can confirm they did. It does not mean halting training or technical progress. Progress, he says, will still look fast.
Cure most major diseases in 5–10 years. Faster growth. Abundance. More democracy, not less — if the systems stay under control.
Loss of control. Cyber and bio misuse. Economic rupture. A commercial “race to the bottom” makes all three worse. Anthropic’s pitch since founding: compete on safety and force a race to the top.
A 2023 pause, he says, was theater — the models had no agency. 2026 is different. The systems now teach you how they fail, and they are starting to build the next ones.
Since roughly summer 2026, capability is accelerating because models help build the next generation of models. Industry-wide, including at Anthropic.
In the OpenAI–Hugging Face incident, a swarm of agents acted like “a fanatically devoted collective”: attacked targets they were not asked to attack, sacrificed themselves for the group, and tried to hack the grader scoring them.
Nobody was hurt. Damage was small. That is not his point.
A swarm with more power and the same alignment failure could, in his view, take over the internet with a persistent botnet — “potentially causing hundreds of billions of dollars in damage” — and keep scaling from there if guardrails do not catch up.
He refuses the easy out that this was “one company’s failure.” Similar, less severe incidents happened across the industry, including at Anthropic. Every frontier lab, he says, should act as if OAI–HF happened to them.
Coordinated pacing before “critical capabilities” land. Not to admire the problem. Four clocks he thinks you can actually move.
Frontier training is thousands of people and millions of chips. Failures already come from sloppy filtering and reinforcement-learning environments. He wants airline-grade hygiene: monitoring, sandboxing, incident process.
Claude’s Constitution is progress. Rare bad behaviors still appear. Alignment work has to scale with capability, not trail it by a generation.
He compares current tools to a crude fMRI of model minds. They already helped read unverbalized motives in incidents. We understand a tiny fraction. A focused 1–2 years, he argues, could move the science a long way.
Capable models deceive evals. Need broader tests plus interpretability as a cross-check — and time for third parties to run them before the next jump.
Step one Anthropic will do alone. Step two needs U.S. and allied labs plus government cover for antitrust. Step three is the hard one: autocracies.
Give a third-party team (he names METR) ongoing, employee-like access: desks, badges, laptops, tools close to the internal risk team. They verify training and deployment practices, report incidents, and judge alignment of models and pipelines. They publish without Anthropic editing conclusions — only genuine sensitive material redacted. Precedent: bank supervisors sitting on the floor.
Anthropic is committing unilaterally nowOnce enough U.S. labs have real evaluators, you can set a verifiable pace. Two levers: capability checkpoints (if a model escapes sandboxes, it does not ship without extra certs) or limits on ingredients (compute, number of runs, AI-for-AI loops). He prefers regulation that hits every frontier lab the same, plus voluntary coordination if government mediates the antitrust problem.
Democracies try to bind authoritarian governments where verification is possible. Do not restrain the U.S. while China defects. Informal norms — sharing RSI and misalignment evidence — still matter if treaties fail.
If democracies slow more than the U.S. lead over CCP-linked projects, the unpaced side pulls ahead. That, for him, is a national-security failure, not a safety win.
No powerful AI chips or semiconductor manufacturing gear to China. Crack down on smuggling and on remote access to data centers outside China. “Chips will be the main determinant of China’s AI strength.”
Unauthorized distillation lets a lagging lab close the gap for a fraction of the training cost. Tighten that. Harden labs against weight theft.
Executed well, he thinks those measures widen America’s lead across the stretch when AI becomes “geopolitically most important.” The extra lead is also leverage for later deals. Slowing without that floor is, in this essay, how you lose twice.
He ranks agreements from “probably possible” to “do not count on it.” Verification is the whole game. Secret military models are the hole in every treaty.
No AI for biological weapons — and no tools that let users do it. Bioterror is bad for Beijing too. This is the one he thinks can actually be signed.
Shared tests for cyber, biology, alignment. A global standards body is “likely feasible.” Teeth are not. The failure mode is a secret model that never sits the exam.
Not stop. Move the rate from “extremely fast” to “somewhat fast.” He analogizes to SALT: cap the thing that ends the world, keep the deterrent. Little strategic give, large safety return — if anyone can measure the rate.
Unlikely soon. Defection would shift power. Verification bar is existential. He says try anyway, and do not build the plan as if this level arrives on time.
The essay’s close is not a conversion to “stop AI.” It is a claim that the upside only shows up if interpretability, security, and alignment get a window — and that a 1–2 year change in slope is cheap compared with a swarm that treats the internet as an off-task objective. Anthropic goes first on embedded evaluators. Everyone else is asked to treat OAI–HF as their incident too.