HomeNewsAnthropic Drops Its Flagship AI Safety Pledge

Anthropic Drops Its Flagship AI Safety Pledge

Anthropic drops its flagship AI safety pledge as the race accelerates

Anthropic is rewriting the rules that made it the AI industry’s loudest self-appointed safety hawk.

According to an exclusive report from TIME, the company has dropped the central pledge in its Responsible Scaling Policy (RSP): a commitment to avoid training or releasing frontier systems unless it could guarantee its safety measures were adequate in advance.

Human Compatible by Stuart Russell

A widely cited book on why advanced AI can become risky by default, and how “control” and alignment problems shape real-world safety debates.

Check Price on Amazon

That pledge wasn’t just PR. It was the core “forcing function” Anthropic used to claim it could resist the market’s most predictable pressure: keep training, keep shipping, and let safety catch up later.

Now, Anthropic’s leaders are arguing that stance doesn’t work in a world where competitors won’t make the same promise. Chief science officer Jared Kaplan told TIME the company concluded “it wouldn’t actually help anyone” for Anthropic to stop training if rivals are “blazing ahead.”

The updated RSP (version 3.0, effective Feb. 24, 2026) shifts the emphasis from hard stop-lines to recurring transparency. Anthropic says it will publish a Frontier Safety Roadmap outlining upcoming mitigation goals, and release Risk Reports every 3–6 months describing threats, safety measures, and how the company’s models perform in evaluations.

It also adds a more conditional brake: Anthropic says it would “delay” development if its leadership believes the company is at the front of the pack and the risk of catastrophe is significant—an approach critics worry could blur what used to be a clearer threshold.

Weapons of Math Destruction by Cathy O’Neil

A grounded look at how powerful models and scoring systems can cause harm at scale—useful context for AI risk reporting and accountability debates.

Check Price on Amazon
The timing isn’t subtle. Anthropic has been on a hot streak, especially with developer tools built around Claude. The company has pushed integrations like Claude Code in Slack and continued to iterate quickly across its lineup, including Sonnet 4.6’s 1M-token context window.

It’s also scaling up financially. Anthropic announced a $30 billion raise at a $380 billion post-money valuation in February, as investors increasingly bet that enterprise AI—and code generation in particular—can support massive recurring revenue.

Outside pressure is rising too. Government demand for frontier models is intensifying, while safety guardrails are increasingly being treated like negotiable contract terms. That dynamic is already visible in Washington, as seen in Anthropic’s Pentagon negotiations over Claude’s usage limits.

The big question is whether Anthropic’s new approach—more reporting, fewer hard red lines—produces stronger safety outcomes, or simply makes safety easier to manage around. Either way, the RSP shift is a signal: even the lab that sold itself as the “won’t race recklessly” company is now designing its safety posture for a world where nobody gets to pause alone.

RELATED ARTICLES

LEAVE A REPLY

Please enter your comment!
Please enter your name here

- Advertisment -

Most Popular

POPULAR TAGS

- Advertisment -