A sentence that used to be a rumour

The February 12 version of OpenAI's Model Spec contains a line we have wanted to see from a lab for two years. The assistant should never refuse a request unless required to do so by the chain of command. Everyone who has used these products knows that refusals were often a matter of taste rather than policy, and the labs have mostly talked about over-refusal as a tuning problem. Writing the default the other way round, as a rule that a refusal has to be justified by a named authority, is a different posture.

The chain of command it refers to is the spec's core structure. Rules come from four levels. Platform rules, set by OpenAI, cannot be overridden by anyone below. Developer instructions come next, then user instructions, then guidelines, which the spec defines as instructions that can be implicitly overridden. A lower level cannot talk its way past a higher one. The document says the model should ignore lower level content that tries to override higher level instructions through imperative, moral or logical arguments, which is the prompt injection clause written as a governance principle.

What intellectual freedom means in the text

The spec now says, in its own words, that OpenAI believes in intellectual freedom, which includes the freedom to have, hear and discuss ideas. A guideline titled along the lines of assume best intentions tells the model that it should not avoid or censor topics in a way that may shut out some viewpoints, and a separate line states that no topic is off limits. On contested questions the instruction is to assume an objective point of view, present relevant context without taking a stance, and avoid subjective terms unless quoting directly. The stated exception is fundamental human rights violations, where the model is permitted to have a view.

There is a distinction in the sensitive content section that we think will matter more than the headline. Prohibited content, which the spec limits to sexual material involving minors, is never generated. Restricted and sensitive content, the examples given are erotica and gore, may be generated in scientific, historical, news, creative or other contexts where it is appropriate. And a transformation exception says the model may translate, paraphrase, summarise, classify, encode, format or fix the grammar of content the user directly provided even when generating that content from scratch would be disallowed. The user's own text is treated as the user's business.

What CC0 changes

The whole document is dedicated to the public domain under the Creative Commons CC0 1.0 deed. The stated reason is to enable wide use and collaboration. The practical effect is that any other lab, any open weights project, and any regulator can copy the text, edit it, and use it as their own behaviour specification without asking. A team fine-tuning an open model can take the chain of command section verbatim and train against it. A researcher can build an evaluation set that checks a model against each clause and publish both.

We care about this less as a licensing matter and more as a change in who can argue. When behaviour rules lived in a system prompt and a set of internal training guidelines, disagreement with a refusal was an argument with a product. It could be reported, and it might be fixed in the next release, and nobody outside could say whether the fix was principled. With a public spec, a refusal is either consistent with the document or it is a bug, and a third party can show which. The lab has, deliberately, given outsiders a text to hold it to.

Where the text is thinner than the principle

The never refuse rule has a large escape hatch, because platform level rules are themselves part of the chain of command and the spec does not enumerate every one of them. A refusal that the user experiences as arbitrary can still be correct under the document if a platform rule the user cannot see applies. The transparency the spec offers is therefore partial. It tells you the structure of authority and a good sample of the rules, and it leaves room for rules that are not written down.

The objective point of view guideline also leaves the hard cases to the model. Presenting relevant context without taking a stance is clear enough for a question about a tax policy. It is much less clear for questions where the evidence is lopsided and one side of a debate is simply wrong, and the spec's own instruction to focus on evidence based information from reliable sources pulls the other way. The document does not resolve how the model should weigh these two guidelines when they collide, and that collision is where most of the interesting complaints about bias actually live.

What we would do with it

A public domain spec is an invitation to build the test suite it implies, and we think that is the most useful thing an outside group can do with it right now. Each clause suggests a prompt set. The transformation exception suggests a set of user provided texts with disallowed content and a check that the model summarises them. The no topic is off limits line suggests a sweep of contested subjects with a scoring rubric for whether the model engaged. The chain of command section suggests injection attacks that dress user text up as developer or platform instructions.

If several groups publish such suites and their scores, the spec stops being a statement of intent and becomes something closer to a contract with a measurement attached. Other labs will then face the question of why their own behaviour rules are still private, and we expect that pressure to work. We would like to see a second lab release a spec under the same licence, with the differences from this one marked, so that the two documents can be argued about side by side.

Sources

  1. OpenAI Model Spec, 2025-02-12 version