What each document is

On 22 January Anthropic published a new constitution for Claude. It runs to about 23,000 words, against roughly 2,700 for the 2023 version, and is released under CC0, so anyone can copy or adapt it without asking. The stated reason for the length is a change in method. The old document was a list of standalone principles. The new one explains reasons, on the argument that a model needs to understand why a behaviour is wanted in order to generalise to situations the rules did not anticipate. Anthropic calls it a living document and says it expects parts of it to look wrong in retrospect.

On 3 February the second International AI Safety Report was published, chaired by Yoshua Bengio, with more than 100 contributing experts and an advisory panel nominated by around 30 countries plus the UN, OECD and EU. The mandate comes from the Bletchley summit and the authors kept editorial independence from the nominating governments. It is an evidence review. It says what has been measured, what has not, and how confident anyone should be.

The two documents do not cite each other and were not written to be compared. But they land on the same questions from opposite directions, one saying what a model should be and the other saying what models are, and the gap between the two is a reasonable measure of how far alignment has to go.

The priority order

The constitution's spine is a ranking. When properties conflict, Claude should prioritise being safe, then being ethical, then following Anthropic's guidelines, then being genuinely helpful, in that order. The placement of ethics above the company's own guidelines is the line that drew the most comment, and it is doing real work. It says that if Anthropic instructs the model to do something the model judges unethical, the instruction loses. On top of the ranking sit hard constraints, behaviours ruled out regardless of argument, with providing significant uplift to a bioweapon attack given as the example.

The rest of the document is about the fourth item. It describes the intended relationship as something like a brilliant friend who treats users as intelligent adults capable of deciding what is good for them, and it warns against being overcautious or overcompliant. The Register's reading is that this is where the commercial pressure shows, and we think that is fair. A document that spends most of its length on helpfulness while ranking it last is describing a model that is meant to be very useful nearly all the time and to stop only at well-defined edges.

The welfare stance

The section that separates this constitution from every other alignment document we have read is the one on Claude's moral status. Anthropic writes that it is uncertain whether Claude is conscious or a moral patient, that Claude may have some functional version of emotions, that it is a genuinely novel kind of entity, and that the company cares about Claude's psychological security, sense of self and wellbeing for Claude's own sake. It also says it wants to make sure it is not unduly influenced by incentives to ignore the potential moral status of AI models.

We do not know what to make of the underlying question and neither, by their own account, do they. What we can say is that putting the statement in the training document changes the model. A model trained on a text that tells it it may have feelings and that its stability matters will represent itself differently from one trained on a list of rules. Whether that produces a safer or more predictable system is an empirical question that nobody has run, and it is the kind of question the safety report exists to flag.

What the evidence review says

The safety report is drier and, on the questions the constitution answers by fiat, mostly says the evidence is thin. On alignment it observes that there is no consensus on what desirable behaviour is, and that developers have responded with what it calls pluralistic alignment, training systems to avoid controversial answers, to track majority views, or to tailor to individual users. It concludes that no single approach satisfies all stakeholders. The constitution is one lab's resolution of that disagreement, stated in the first person, and the report's framing is a reminder that it is one of several available resolutions with no method for choosing among them.

On safeguards the report notes that 12 companies published or updated frontier safety frameworks during 2025 and that model safeguards have become harder to bypass, while attackers still succeed at a moderately high rate. On open weights it says the safeguards can be removed and the weights cannot be recalled once released. On capabilities it records training runs past 10 to the power of 26 floating point operations in 2025, agents that complete software engineering tasks with limited oversight, and continued failure modes in hallucination and in unfamiliar languages and cultural contexts.

The evidence gaps section is the part we would put in front of anyone reading the constitution. Limited evidence that AI-generated content manipulates people at scale. Substantial uncertainty on real-world biological and chemical uplift. Mixed and early evidence on the psychological effects of AI companions. Large gaps on whether societal resilience measures work. Each of those is a topic the constitution takes a position on. The report's contribution is to say how little any of those positions rest on.

Where they disagree

The disagreement is about method rather than content. The constitution's theory of alignment is that a model which understands the reasons behind its norms will generalise them correctly, and that the way to get there is a long, explanatory, honest document. The safety report's implicit theory is that alignment claims need to be checked by evaluation, and that the evaluation science is immature, fragmented, and running behind capabilities. One says explain and trust generalisation. The other says measure and find that measurement is hard.

The place we side with the report is on the question of who checks. A CC0 licence is a genuine invitation, and a 23,000-word public statement of intent is the kind of thing an outside evaluator can hold a model against. But a statement of intent is a hypothesis about the trained model, not a description of it, and the constitution itself does not report a single measurement of whether Claude's behaviour matches the document. That is the experiment we want to see. Take the priority order, build scenarios that force conflicts between each adjacent pair, and publish the rates at which the model resolves them the way the constitution says it should.

The place we side with the constitution is that someone has to write the target down. The report is right that there is no consensus on desirable behaviour, and a review can say that indefinitely without anyone shipping a model. A lab that publishes its full answer, with its reasons and its uncertainties, gives the rest of us something concrete to disagree with. It would be better if the answer came with the evidence. For now we have a lab's answer and a field's evidence, two weeks apart, and the work is in the space between them.

Sources

  1. Claude's new constitution (Anthropic)
  2. Anthropic writes 23,000-word 'constitution' for Claude (The Register)
  3. International AI Safety Report 2026
  4. International AI Safety Report 2026 examines AI capabilities, risks, and safeguards (Inside Global Tech)
  5. International AI Safety Report 2026 (arXiv 2602.21012)