Jan 22, 2026 · View original article

Anthropic Publishes a New Constitution for Claude, Prioritising Safety Over Helpfulness

On 22 January 2026 Anthropic released a rewritten constitution for its Claude models, ordering priorities as safe, ethical, guideline-compliant and helpful, and licensing the text under CC0.

On 22 January 2026 Anthropic published a substantially rewritten constitution for its Claude models, the natural-language document that shapes how the models are trained to behave. The company said the text is written primarily for Claude itself and departs from the earlier approach of listing standalone principles. Instead it tries to explain the reasoning behind expected behaviour, on the theory that a model which understands why a rule exists will generalise better to situations the rule never anticipated.

The document sets an explicit ordering of priorities: Claude should be broadly safe, broadly ethical, compliant with Anthropic's own guidelines, and genuinely helpful, in that sequence when the goals conflict. Its main sections cover helpfulness and user benefit, Anthropic's specific guidelines, ethical standards, safety and human oversight, and a closing discussion of uncertainty about the model's own nature. Anthropic states that the constitution feeds into several stages of training, including the generation of synthetic data and the ranking of candidate responses. The full text is released under a Creative Commons CC0 1.0 dedication, which allows anyone to reuse or adapt it without restriction.

Anthropic's stated rationale for publishing is accountability: if the intended behaviour is public, users and researchers can distinguish intended from unintended conduct and give feedback. The company also noted that this matters more as models take on agentic tasks with greater real-world influence.

The release is best understood as an artefact of AI governance rather than a product launch. Frontier labs are under growing pressure, from California's new transparency law, from the EU AI Act's obligations on general-purpose model providers and from enterprise customers, to show how model behaviour is specified and controlled. A published constitution is one answer to that demand, and the CC0 licence invites others to build on it. It is also a reminder of the limits of such documents: a constitution describes intent, not guaranteed behaviour, and the ordering of safety above helpfulness is itself a design decision that some customers will welcome and others will find restrictive when the model declines a request.

For organisations deploying Claude, the practical value lies in having a citable description of intended behaviour that can be referenced in their own risk assessments and vendor due diligence.

What it means for leaders

  • Add published behaviour specifications to vendor due diligence. A model's constitution or spec is evidence for the "Govern" and "Map" functions of the NIST AI RMF and for supplier evaluation under ISO/IEC 42001.
  • Do not mistake a constitution for a control. It states intent; your own testing, monitoring and human-oversight measures remain necessary, particularly for agentic deployments.
  • Align your own AI policy with the vendor's priority ordering. If safety sits above helpfulness, workflows should anticipate refusals and route them to humans rather than treating them as outages.
  • Use the CC0 text as a template. Organisations drafting internal AI principles can adapt the structure while stating their own obligations and escalation paths.
  • Expect regulators to ask for equivalent documentation. Transparency requirements for general-purpose models in the EU and for frontier developers in California point in the same direction.

Comments

No comments yet. Be the first to comment.