
Structured Outputs Explained: JSON Mode vs Strict Schemas vs Constrained Decoding
Constrained decoding now guarantees an LLM returns schema-valid JSON, but valid JSON is not correct JSON. How it works under the hood, what the latest research shows it cannot fix, and the three-layer pattern that holds up in production, with free open-source options.
You asked the model for JSON. It gave you JSON wrapped in a markdown code fence, with a friendly "Sure! Here's the data:" on top, and one field called quantity instead of qty. Your parser threw, your retry loop fired three times, and the workflow timed out. Every team that ships an LLM feature hits this wall, and most of them fix it by bolting on more prompt instructions. That's the wrong layer. In 2026 the fix lives in the decoder, and understanding exactly what it guarantees, and what it quietly doesn't, is the difference between a demo and a product.
JSON Mode vs Structured Outputs vs Tool Calling: Three Different Promises
These three features get used interchangeably, but they promise different things.
- JSON mode only guarantees the output parses as JSON. It says nothing about which keys appear or what types they hold. OpenAI's own Structured Outputs guide is explicit that only Structured Outputs ensures schema adherence; JSON mode does not.
- Structured outputs (strict schemas) guarantee the response matches a JSON Schema you supply: required keys present, types correct, no extra fields. In OpenAI's strict mode every property must be listed as required and additionalProperties must be false; you express an optional field as a union with null, not by leaving it out.
- Strict tool calling applies the same guarantee to the arguments of a function call. Anthropic's structured outputs docs split this into two capabilities: output_config.format for the response itself, and strict: true on a tool definition for its inputs.
| Feature | Guarantees valid JSON | Guarantees your schema | Where you find it |
|---|---|---|---|
| JSON mode | Yes | No | OpenAI (older models too), most local servers |
| Structured outputs (strict) | Yes | Yes, for supported keywords | OpenAI, Claude, Gemini, Ollama, vLLM |
| Strict tool calling | Yes (arguments) | Yes, for tool inputs | OpenAI, Claude |
If your code does JSON.parse() and then reads specific fields, JSON mode is not enough. You want a strict schema.
How Constrained Decoding Actually Works
A language model doesn't write JSON. At every step it produces a probability for every token in its vocabulary, roughly 100,000 to 200,000 options, and a sampler picks one. Prompting for JSON just nudges those probabilities. Constrained decoding changes the rules of the pick.
Before generation starts, the engine compiles your JSON Schema into a grammar, effectively a state machine that knows every legal continuation of a partial output. At each step it looks at what's been generated so far, works out which vocabulary tokens would keep the output a valid prefix of *some* schema-conforming document, and sets the probability of every other token to zero. If the output so far is {"qty": and the schema says qty is an integer, the tokens for 3 or 12 survive; "three", Sure, and a markdown backtick are masked out. The model still chooses among the allowed tokens, so it's still the model's judgement, just fenced in.
The hard engineering is doing that mask fast. Tokens don't line up with JSON characters (one token might be ":" or },{), so the engine has to reason about multi-character tokens against a stack-based parser on every step. XGrammar, the Apache-2.0 engine that is now the default structured-generation backend in vLLM, SGLang, TensorRT-LLM and MLC-LLM, precomputes most of that work so the per-token overhead is close to zero. That's also why both OpenAI and Anthropic warn that the *first* request with a new schema is slower: that's grammar compilation. Anthropic caches compiled grammars for 24 hours from last use, and changing only a tool's name or description doesn't invalidate the cache.
What the Numbers Say: Structure Is Solved
A September 2026 preprint, Constrained Decoding Eliminates Structural Failures in Small LLMs but Reveals a Scale-Dependent Semantic Gap, tested five small open models (0.6B to 4B parameters) on 14 structured-output tasks, with and without constrained decoding through both Outlines and XGrammar. Prompt-only, the Llama 3.2 1B and 3B models produced schema-valid output 78.6% of the time. With either constraint engine, every model hit 100%.

The same study measured the cost. XGrammar added 4 to 8 ms of schema compilation and a 1.6 to 3.7% throughput reduction. Outlines averaged 2.0 to 4.9 seconds of compilation per model, peaking at 19.5 seconds on the most complex schema, and ran 9 to 13% slower on four of the five models. It's a single-author preprint with 210 evaluations, so treat the exact figures as directional, but the shape matches what you'll see in your own logs: structural failures go to zero, and the engine you pick decides whether that's free.
Valid JSON Is Not Correct JSON
Here's the part the marketing pages skip. The grammar checks shape, not meaning. The same paper describes a "hollow rescue": Llama 1B under XGrammar scored perfectly on a multi-step function-calling task, with every call and every parameter present, while the email body it produced was just a numbered list with no content. Every structural metric passed. The output was useless.
A separate May 2026 benchmark, OrderBench, makes the risk concrete for agents that turn user intent into API calls. Across 2,400 calls to four open models acting as restaurant-ordering agents, the strongest model hit 100% schema validity in both prompt-only and JSON-schema modes, yet semantic success stayed near 80%. Weaker models made schema-valid *unsafe acceptances*, orders that should have been refused, at double-digit rates. The author's conclusion is the right mental model: structured output is a necessary interface layer, not a substitute for domain verification.
There are also mechanical gaps worth knowing:
- Not every JSON Schema keyword is enforced. Claude's docs list minimum, maximum, multipleOf, minLength and maxLength as unsupported in the grammar. The official SDKs strip them before sending, move the constraint into the field description, and re-validate the response against your original schema afterwards. A qty of 300 against maximum: 20 is still schema-valid as far as the decoder is concerned.
- Truncation breaks the guarantee. If generation hits max_tokens mid-object, you get half a document. Research like TruncProof exists precisely because of this. Check the stop reason before you parse.
- Refusals are a separate channel. OpenAI returns a model refusal as a distinct refusal field, not as a schema-shaped object, so your code needs a branch for it.
- Field order matters. Generation is left to right. If your schema puts answer before reasoning, the model commits to the answer before it has written its reasoning. Put working fields first, verdict fields last.
A Production Pattern That Holds Up
Treat the decoder as the first of three layers, not the only one.
- Layer 1: constrain the shape. Use strict structured outputs on hosted APIs, or a grammar engine when you self-host. Keep schemas flat, use enum wherever the valid values are known, and give every field a precise description, since that description is the only semantic guidance the model gets inside the fence.
- Layer 2: validate the rules. Parse into Pydantic or Zod with the business constraints the grammar can't express: ranges, cross-field rules ("end date after start date"), IDs that must exist in your database. Instructor (MIT, around 14k GitHub stars) wraps this loop and feeds validation errors back to the model for a targeted retry.
- Layer 3: fail closed on consequences. Anything that spends money, sends a message, or writes to a system of record gets checked against deterministic rules before it executes, and ambiguous cases go to a human. This is the same principle behind the approval gates in our guide to agent guardrails and human-in-the-loop oversight and the tool-permission limits in the n8n prompt injection breakdown.
Then measure semantic accuracy separately from schema validity. A dashboard that says "100% valid JSON" is measuring the part you already solved; the eval approaches in Evaluating AI Agents in Production are how you track the part you haven't.
Free and Open-Source Options if You Self-Host
You don't need a paid API to get guaranteed structure. The open stack is mature:
- Ollama: pass a JSON Schema in the format parameter (announcement and examples). Ollama's own tips are to define the schema with Pydantic or Zod, say "return as JSON" in the prompt, and set temperature to 0. The easiest start if you already run a local model with Open WebUI.
- vLLM: its structured outputs feature supports JSON Schema, regex, choice lists and full context-free grammars through the OpenAI-compatible response_format field, with xgrammar or guidance as backends. The old guided_json parameters were removed in v0.12.0, so update older tutorials' code.
- XGrammar and llguidance: the grammar engines underneath, usable directly if you're building your own inference path.
- Outlines (Apache-2.0, around 16k stars): the Python library that popularised constrained generation, good for experimenting with Hugging Face models locally. Budget for its compile time on complex schemas.
If you want to compare engines on your own schemas, the JSONSchemaBench benchmark from EPFL and Microsoft researchers ships 10,000 real-world schemas and was built to test exactly these frameworks, including OpenAI and Gemini.
Where to Go Deeper
Structured outputs are one of those features that look like a checkbox and turn out to be an architecture decision. Getting the decoder right takes an afternoon. Designing the schema so the model reasons before it commits, writing the validation layer that catches what the grammar can't, and building evals that measure whether the content is actually right is the work that decides whether an AI feature survives real users.
That's the ground the AI Product Development Bootcamp covers end to end: turning LLM calls into typed, validated interfaces, wiring them into RAG and agent pipelines, and setting up the evaluation and failure handling that lets you ship with confidence instead of retry loops. If this post was the "why," the bootcamp is the build.
Go deeper
AI Product Development Bootcamp
Want the full AI Product Development Bootcamp course, not just this post?
Join the waitlist and get a 20% launch discount the moment we open checkout. No payment now.
Taught by Aditya Jha · 40+ AI products shipped for real clients. No spam, unsubscribe any time.
Or join our free community for AI tips while you wait
One email a week: the AI tools, tactics, and course drops actually worth your time. No spam, unsubscribe anytime.
Questions about this or which course fits? Email academy@aibootstrapper.com and we'll answer it directly, not with a support ticket.