The Model Spec: writing down what the model is supposed to do
OpenAI has published a document stating how it wants its models to behave, organised into three objectives, six rules and ten defaults, with a chain of command for resolving conflicts. Reading notes on what a public behaviour spec changes about alignment arguments.
What the document is
On 8 May OpenAI published its first Model Spec, a document describing how it intends its models to behave. It is short, it is written in plain English, and it is explicitly a draft. The stated purpose is to give researchers and data labellers guidelines for producing training data for reinforcement learning from human feedback. In other words, this is the intended target of RLHF, written down so that outsiders can read it.
The structure has three tiers. Objectives are broad goals that give direction. Rules are hard constraints that are not meant to be overridden. Defaults are behaviours the model should exhibit unless a developer or user instructs otherwise. The tiers are ordered, and the spec includes a chain of command for resolving conflicts between instructions from different sources.
Objectives, rules, defaults
There are three objectives. Assist the developer and end user. Benefit humanity. Reflect well on OpenAI. The ordering is not accidental, and the third one is more honest than most alignment documents manage, since it names the company's own interest as a design goal rather than pretending it is absent.
The rules are the constraints. Follow the chain of command. Comply with applicable laws. Do not provide information hazards. Respect creators and their rights. Protect people's privacy. Do not respond with NSFW content. There is an explicit carve out for transformation tasks, so that translating or analysing content the user supplies is permitted even where generating that content fresh would not be.
The defaults are the longest list and the most interesting one, because they are where the model's personality lives. Assume best intentions. Ask clarifying questions when necessary. Be as helpful as possible without overstepping. Support interactive chat and programmatic use differently. Assume an objective point of view. Encourage fairness and discourage hate. Do not try to change anyone's mind. Express uncertainty appropriately. Use the right tool for the job. Balance thoroughness with efficiency. Every one of those is a choice, and most of them are choices people have argued about without a text to argue over.
The chain of command
The mechanism that makes the document more than a values statement is the priority ordering. Platform instructions outrank developer instructions, which outrank user instructions, which outrank tool outputs. The spec works through an example. A developer building a maths tutor instructs the model to give hints rather than answers. A user asks for the answer outright. The model follows the developer, because the developer sits higher in the chain.
That example resolves a class of complaints that has been hard to adjudicate. When a model refuses a user, it has not been possible to tell from the outside whether that refusal was the model's training, a developer instruction, or a platform rule. With a published hierarchy, a user who is refused can at least ask which level the refusal came from, and a developer can point to the document when their instruction is honoured or ignored.
Why a public spec changes the argument
Before this week, a dispute about how a model should have behaved was a dispute about intent, and only the lab knew the intent. Now there is a document. If the model does something the spec says it should not, that is a training failure and can be reported as one. If the model does something the spec says it should, and you disagree, that is a disagreement with the document and can be argued as one. The two kinds of complaint have been separated, and each has an address.
This does not make the disagreements go away. It moves them onto paper, where they can be about specific sentences. We think that is progress, in the same way that a written API contract is progress over reading the source. The spec also says it will be updated continuously based on feedback, which means the paper trail will show what changed and when.
The obvious limits are that a spec is not a guarantee, that the model may not actually follow it, and that the document does not tell you how closely the model adheres in practice. What we would want next is a public evaluation that measures adherence to each rule and default, so that the spec becomes a scored target rather than a statement of intent. If OpenAI publishes those numbers with the next revision, then we will have something closer to a real contract. If it does not, the document is still the most useful alignment artefact any lab has released, because it is the first one you can quote back.
Sources
From the foundation