Practices/Exemplar policies

institutional policy

Anthropic's Responsible Scaling Policy

Anthropic. Version 3.4, effective 8 July 2026 (page last updated 14 August 2026).

Summary

Anthropic publishes a voluntary framework for scaling frontier models. The policy sets AI Safety Levels and capability thresholds, including for CBRN-related misuse and for AI that can accelerate research. Version 3.0, announced 24 February 2026, separated what Anthropic says it will do on its own from what it says the industry should do together, and it added a Frontier Safety Roadmap and periodic Risk Reports. Anthropic is cited here as a source and as an example of a public scaling policy. That citation does not mean Anthropic endorses, partners with, or is affiliated with AI Concerns.

The principles it rests on

In our words.

  1. Proportional safeguards: stronger models face stronger deployment and security standards.
  2. Named thresholds: capability levels trigger named safety levels, rather than leaving the trigger unstated.
  3. Unilateral and shared work: some commitments are the company's; some risks, the authors say, need the industry.
  4. Public reporting: Risk Reports are part of the policy, not an afterthought.
  5. Stated exclusions: the authors name what a given standard does not cover.

Scope

Anthropic's own training and deployment of frontier models, with public thresholds and safety levels. The ASL-3 Security Standard, as described by the authors, does not cover sophisticated insiders or state-compromised insiders.

Why it is exemplary

It is a public, versioned commitment with named thresholds, not only a values page. It says what will trigger a higher bar. It also says what a given standard does not cover. Readers can disagree with the thresholds and still see how a scaling policy is written.

Limits the authors state themselves

The revised RSP aims to adopt more realistic unilateral commitments that are difficult but still achievable in the current environment, while continuing to comprehensively map the risks we believe the full industry needs to address multilaterally.

Anthropic, announcing Responsible Scaling Policy Version 3.0, 24 February 2026.

Concerns in our register it speaks to

Frameworks it relates to

  • Anthropic AI Safety Levels (ASL), as named by the authors
  • Frontier Safety Roadmap and Risk Reports (as named by the authors)

How a reader might use it

  • A school

    Use it to show what a public threshold looks like, then ask who sets it and who can contest it.

  • A health system

    Use it as a model of naming the misuse path (for example CBRN) before a capable system is scaled.

  • A public body

    Use it as a model of a versioned, public scaling policy, including what the authors say they will not cover alone.

Source

Anthropic Responsible Scaling Policy

Anthropic, Responsible Scaling Policy, Version 3.4, effective 8 July 2026. Version 3.0 announced 24 February 2026.

Summarized with attribution. The original is the property of its publisher; read it at the link.