From function calling to structured outputs: making JSON a contract
OpenAI now guarantees that model output matches a JSON schema. That fixes syntax, which was most of the pain. It does nothing for the values inside, which is where the failures have moved.
Fourteen months of asking nicely
On August 6 OpenAI shipped Structured Outputs with a new model, gpt-4o-2024-08-06. On their own evaluation of complex schema following, the new model in strict mode scores 100 percent, while gpt-4-0613 scored under 40 percent. That second number is the one to sit with. For fourteen months a lot of us have been building systems on a model that got the shape of its output wrong more than half the time on hard schemas, and papering over it with retries and regex.
Function calling arrived on June 13, 2023, on the 0613 versions of GPT-4 and GPT-3.5-turbo, alongside a 16k context variant and price cuts. The pitch then, as TechCrunch reported it, was that developers could describe functions and the model would produce the arguments to call them, which let it "more reliably get structured data back". More reliably was accurate. Reliably was not. The model was fine-tuned to prefer JSON, and it usually did, and usually is a terrible word to build a parser around.
What changed under the hood
The new mechanism is constrained decoding. Simon Willison notes it is the same family of technique as jsonformer and the grammar support in llama.cpp. Ted Sanders at OpenAI explained that the first request with any schema is slow because the schema is preprocessed into a context-free grammar, after which the grammar is cached. At each decoding step the sampler only permits tokens that keep the output inside the grammar. A closing brace cannot appear before the required keys are present. An enum field cannot take a value outside the enum.
OpenAI describes this as a two-part effort. The model was trained to follow complicated schemas, and then the engineering layer was added on top to take the last few percent to 100. That ordering matters. Constrained decoding on a model that has no idea what the schema means produces valid JSON full of nonsense. The training got the model close and the grammar removed the residual syntax errors.
Two API surfaces expose it. A strict flag on function definitions, and a json_schema response format for when you want a structured answer without a tool call. There is a Pydantic-based helper in the Python library. The model also got cheaper and bigger: output limit up from 4,096 to 16,384 tokens, input at $2.50 per million and output at $10.00, down 50 and 33 percent respectively.
What the guarantee actually covers
The guarantee is syntactic. The output will parse, every required key will be present, every enum value will be a member of the enum, and every type will match. The docs say this directly and then add the caveat people skip: "Structured Outputs can still contain mistakes." A schema that says a field is an integer between one and ten guarantees you get an integer between one and ten. It says nothing about whether that integer is the right one.
The documented restrictions are the price of getting a grammar out of a schema. Every field must be listed in required, so optional fields are simulated with a null union. Every object must set additionalProperties to false. The root must be an object, not an anyOf. Defaults are unsupported. Supported types are string, number, boolean, integer, object, array, enum and anyOf, with a limited set of string formats and numeric bounds. If you have a large existing schema with optional fields, it will not go through strict mode unchanged.
There is one genuinely new signal. When the model refuses on safety grounds it returns a refusal field rather than trying to fit a refusal into your schema. Before this, a refusal would either break the parser or, worse, come back as a schema-valid object with an apology in a string field that downstream code treated as data.
Where the failures moved
Here is the pattern we expect teams to hit now that the parser never throws. Ask for a list of the people mentioned in an email with a required role field taken from an enum of manager, engineer and customer. The model finds a name, the role is unclear, and the grammar will not allow an empty value, so it picks one. The output validates. The role is wrong. Before strict mode the model might have written "unknown" and your parser would have caught it as an invalid enum. Now it is silently confident.
The same thing happens with required fields the source does not contain. A schema demanding a date for every invoice will get a date for every invoice. Constrained decoding cannot refuse to fill a slot. The fix is to design the schema so that absence is representable, which usually means more null unions than feels natural, and then to test the values, not just the shape.
A worked example of what testing should look like. Take a hundred documents with hand-labelled extractions. Run the extraction, then compare field by field. Report three numbers: the parse rate, which is now 100 percent and uninteresting, the field-level accuracy, and the rate at which the model invented a value for a field the document did not contain. The third number is the one that tells you whether the contract is being honoured in the sense that matters.
Why this matters for agents
Tool use is a loop of asking the model for arguments, running the tool, and feeding the result back. When one call in twenty produced malformed arguments, a ten-step agent failed roughly forty percent of the time on syntax alone before any reasoning error had a chance to occur. Removing that failure class is what makes multi-step tool use something you can put a service level on. That is the quiet reason this release is more important than the model refresh that came with it.
What it does not do is make the arguments correct. An agent that calls delete_file with a valid path is still an agent that may have picked the wrong path. The next layer of reliability has to come from validating semantics at the tool boundary, and we would like to see people publish the kinds of checks they put there. Our own guess is that the winning pattern is a schema for shape plus a small deterministic validator for meaning, and that the second part is where the remaining engineering lives.
Sources
From the foundation