Anthropic

Anthropic's RSP 3.0 Drops Its Automatic Pause Pledge

On Feb 24, 2026 Anthropic shipped RSP 3.0, dropping its pledge to automatically pause work on dangerous models and recasting safety around competitors' actions. What changed, why, and the reactions.

Anthropic's RSP 3.0 Drops Its Automatic Pause Pledge — article cover
On this page6 SECTIONS
  1. What the Old Pledge Was
  2. What Changed in RSP 3.0
  3. Anthropic’s Reasoning
  4. Reactions
  5. What It Means for Developers and Enterprises
  6. Sources

On February 24, 2026, Anthropic published version 3.0 of its Responsible Scaling Policy, and by the next day the story was everywhere: TIME ran it as an exclusive under the headline “Anthropic Drops Flagship Safety Pledge,” CNN reported that the company was ditching its core safety promise, and the Wall Street Journal carried it in print. A lab whose founding identity was safety-first had just rewritten its most famous safety rule.

The change boils down to one sentence. Anthropic will no longer automatically pause development of a model that could be considered dangerous. The new policy states that if one developer pauses to implement safety measures while rivals keep training and deploying without strong mitigations, the result could be “a world that is less safe” — “the developers with the weakest protections would set the pace, and responsible developers would lose their ability to do safety research.” Anthropic says this language was approved unanimously by CEO Dario Amodei and the board.

What the Old Pledge Was

The RSP arrived in September 2023 as a voluntary, if-then framework: when a model’s capabilities cross a threshold (labeled with ASL levels), matching safeguards must be in place first. Its influence ran well beyond one company. OpenAI and Google DeepMind adopted similar frameworks within months, and its fingerprints are visible in California’s SB 53, New York’s RAISE Act, and the EU AI Act Codes of Practice. Anthropic’s frontier models have operated under ASL-3 safeguards since May 2025.

The load-bearing clause was the commitment to have stringent model security measures ready in advance of training — what TIME called an “ironclad” guarantee. It never actually triggered a pause. But it was the anchor of the entire voluntary-safety argument.

What Changed in RSP 3.0

Three substantive moves.

First, the automatic pause is gone. Under the new text, leadership would consider delaying development only when it simultaneously believes Anthropic is the race leader and that catastrophe risks are significant; the replacement commitment is to “match or surpass” competitors’ safety efforts.

Second, company plans are now separated from industry recommendations: one set of mitigations applies to Anthropic’s own systems, and a separate set is what the company recommends to the field, so its own conditions are no longer quietly externalized into everyone else’s obligations.

Third, a Frontier Safety Roadmap and recurring Risk Reports. The roadmap is a public, nonbinding list of goals with graded progress — the announcement’s own examples include a moonshot R&D program for information security, automated red-teaming that surpasses the contributions of human bug bounties, measures for adherence to model constitutions, AI analysis of centralized records to catch insider threats, and a “regulatory ladder” policy roadmap. Risk Reports follow every three to six months, covering model properties relevant to catastrophic risk and the state of mitigations, with third-party expert reviewers granted unredacted access in defined circumstances.

Anthropic’s Reasoning

Co-founder and chief science officer Jared Kaplan was blunt with TIME: “We felt that it wouldn’t actually help anyone for us to stop training AI models,” and “we didn’t really feel, with the rapid advance of AI, that it made sense for us to make unilateral commitments… if competitors are blazing ahead.”

The announcement adds structural reasons: capability evaluations live in a “zone of ambiguity” where models approach but don’t definitively cross thresholds, which the old mechanism had no answer for; hitting RAND’s recommended SL5 security standard unilaterally is “currently not possible”; and the political climate around regulation has shifted. TIME supplies the commercial backdrop — Anthropic had just raised $30 billion in February at a reported $380 billion valuation, with annualized revenue that TIME says has grown roughly tenfold each year. Mashable, citing Axios, notes the Department of Defense had pressured Anthropic to open up military applications under threat of restrictions.

Reactions

Chris Painter, director of policy at evaluator METR, told TIME the labs are in “triage mode” and that “society is not prepared.” His specific worry is that removing binary thresholds trades tripwires for a slow accumulation of risk — a “frog-boiling” effect. Advocacy groups like the Transparency Coalition treated the move as more evidence that voluntary corporate pledges cannot be relied on and that safeguards must be written into law.

The honest counterpoint is that Anthropic is simultaneously increasing transparency: the revised RSP, the recurring risk reports, and January’s fully public, CC0-licensed Claude constitution all follow a “publish it and let people check” line. What was dropped is the hard commitment to stop; what remains is the soft commitment to disclose.

What It Means for Developers and Enterprises

Three practical consequences. First, procurement and governance: teams that treated vendor self-restraint as a backstop now have that backstop explicitly subordinated to race dynamics — draw your own red lines in contracts and evaluation pipelines. Second, policy: state bills that codify RSP-style thinking (SB 53 in California, the RAISE Act in New York) lose none of their technical reference value, but the argument that voluntary frameworks suffice has been weakened by its own inventor — a gift to the pro-legislation camp. Third, watch the peers: whether OpenAI’s and Google DeepMind’s parallel frameworks get similar revisions is the signal that matters next. If all three loosen in sync, “race to the bottom” stops being a slogan and becomes a description.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

SHAREXEMAIL