Structured Output from LLMs in Python with Pydantic
Turn model replies into validated Python objects: one Pydantic schema, schema-constrained generation where available, and a repair loop everywhere else.
A team asked a model to “reply in JSON” and parsed the result with json.loads. It worked in every test. In production, about one reply in two hundred began with “Sure! Here is the JSON you asked for:”, and the parser crashed. They added a regular expression to strip the preamble. Then a reply arrived with a trailing comma, then one with a missing field, then one with "priority": "high" where the code expected a number.
LLM structured output in Python means getting a model to return data your program can use directly: a validated object with known fields and types, instead of prose you must parse. The standard tool for the Python side is Pydantic, a library that validates data against a class with type annotations. For a full introduction, see Pydantic v2 models and validators.
You need Python 3.13 and Pydantic 2. The complete example runs offline with scripted model replies, so it needs no API key and costs nothing. The provider examples call real APIs, which are billed per token. Model output is non-deterministic, so the same input can produce different values on different runs.
My position: a schema guarantees form, and only your code can check substance. Teams that adopt constrained generation and then stop validating have traded visible crashes for invisible wrong data, which is the worse failure.
Four ways to get structure, from weakest to strongest
| Approach | What it guarantees | What can still go wrong |
|---|---|---|
| Ask for JSON in the prompt | Nothing | Extra prose, broken syntax, missing or renamed fields, wrong types |
| JSON mode | The reply is syntactically valid JSON | Any fields, any types. It may not match your schema at all |
| Schema-constrained generation (Structured Outputs) | The reply matches your JSON Schema | Values can be wrong, invented, or out of your business range |
| Forced tool call with a schema | Arguments shaped by the schema, with strength depending on the provider | The same as above, plus occasional schema misses on some providers |
The third row deserves a short explanation. A model generates one token at a time. With constrained generation, the provider restricts each step to tokens that keep the output valid under your schema. The model cannot emit a stray sentence or a wrong type, because those tokens are never available. OpenAI introduced this as Structured Outputs in August of this year.
Whichever row you use, the same Python code sits on the receiving end: a Pydantic model that validates the result. That is why the model class comes first.
Set up
python --version
python -m venv .venv
source .venv/bin/activate # macOS and Linux
.venv\Scripts\Activate.ps1 # Windows PowerShell
python -m pip install pydantic pytest
That is enough for the offline example and the tests. For the provider sections, also install openai or anthropic, and set the matching API key in an environment variable. Keys never belong in code.
LLM structured output in Python: the complete example
The task is to turn a customer message into a support ticket. Save this as extract.py.
import json
import sys
from collections.abc import Callable
from typing import Literal
from pydantic import BaseModel, ConfigDict, Field, ValidationError, field_validator
OPENAI_MODEL = 'gpt-4o-2024-08-06'
ANTHROPIC_MODEL = 'claude-3-5-sonnet-20241022'
MESSAGE = (
'Hi, I was charged twice for my annual plan this morning. '
'Please refund one payment. You can reach me at dana@example.com. Thanks, Dana'
)
class ExtractionError(Exception):
"""Raised when no valid ticket could be produced."""
class Ticket(BaseModel):
model_config = ConfigDict(extra='forbid')
category: Literal['billing', 'bug', 'account', 'other']
priority: int = Field(description='1 is the lowest urgency and 5 is the highest')
summary: str = Field(description='One sentence of at most 20 words')
customer_email: str | None = Field(description='Email address found in the message, or null')
needs_human: bool = Field(description='True if a person must act on this ticket')
@field_validator('priority')
@classmethod
def priority_in_range(cls, value: int) -> int:
if not 1 <= value <= 5:
raise ValueError('priority must be between 1 and 5')
return value
SYSTEM_PROMPT = (
'Extract a support ticket from the user message. '
'Reply with JSON only, matching this JSON Schema:\n'
+ json.dumps(Ticket.model_json_schema())
)
def check_grounding(ticket: Ticket, source: str) -> list[str]:
"""Return problems where an extracted value does not appear in the source text."""
problems: list[str] = []
email = ticket.customer_email
if email and email.lower() not in source.lower():
problems.append(f'customer_email {email!r} does not appear in the message')
return problems
def extract_with_repair(
generate: Callable[[list[dict]], str], text: str, max_attempts: int = 3
) -> Ticket:
messages = [
{'role': 'system', 'content': SYSTEM_PROMPT},
{'role': 'user', 'content': text},
]
problems: list[str] = []
for attempt in range(1, max_attempts + 1):
raw = generate(messages)
try:
ticket = Ticket.model_validate_json(raw)
except ValidationError as error:
problems = [
f"{'.'.join(str(part) for part in item['loc'])}: {item['msg']}"
for item in error.errors()
]
else:
problems = check_grounding(ticket, text)
if not problems:
return ticket
print(f'attempt {attempt} rejected: {problems}', file=sys.stderr)
messages.append({'role': 'assistant', 'content': raw})
messages.append(
{
'role': 'user',
'content': 'Your reply was rejected:\n- '
+ '\n- '.join(problems)
+ '\nReply again with corrected JSON only.',
}
)
raise ExtractionError(f'no valid ticket after {max_attempts} attempts: {problems}')
def scripted_generate() -> Callable[[list[dict]], str]:
"""Offline stand-in for a model: two flawed replies, then a correct one."""
base = {
'category': 'billing',
'summary': 'Customer was charged twice for the annual plan.',
'needs_human': True,
}
replies = [
json.dumps(base | {'priority': 9, 'customer_email': 'dana@example.com'}),
json.dumps(base | {'priority': 4, 'customer_email': 'dana.smith@example.com'}),
json.dumps(base | {'priority': 4, 'customer_email': 'dana@example.com'}),
]
return lambda messages: replies.pop(0)
def main() -> int:
try:
ticket = extract_with_repair(scripted_generate(), MESSAGE)
except ExtractionError as error:
print(f'extraction failed: {error}', file=sys.stderr)
return 1
print(ticket)
return 0
if __name__ == '__main__':
raise SystemExit(main())
python extract.py
attempt 1 rejected: ['priority: Value error, priority must be between 1 and 5']
attempt 2 rejected: ["customer_email 'dana.smith@example.com' does not appear in the message"]
category='billing' priority=4 summary='Customer was charged twice for the annual plan.' customer_email='dana@example.com' needs_human=True
The scripted model makes two realistic mistakes. First, it returns a priority of 9, which is valid JSON and an integer, and still wrong. Second, it returns an email address that looks plausible and does not exist in the message. Both replies would pass a syntax check. The third reply passes every check, and the function returns a typed Ticket.
What the Ticket class does
- It is the schema.
Ticket.model_json_schema()produces the JSON Schema placed in the prompt. Field descriptions become instructions to the model, so write them carefully. - It is the parser.
Ticket.model_validate_json(raw)turns text into an object or raises aValidationErrorthat lists every problem. - It narrows the choices.
Literal['billing', 'bug', 'account', 'other']allows four categories and nothing else. Always include an escape value such as'other'. Without one, the model must force every message into a category that may not fit. - It holds rules the schema cannot. The
priority_in_rangevalidator runs in Python after parsing.
customer_email: str | None has no default, so the field is required but may be null. That is deliberate. A model that must write null explicitly is making a decision, while a model that may omit a field can simply forget it.
The repair loop
extract_with_repair takes any function that maps messages to text. When validation fails, it appends the model’s own reply and a message listing the exact problems, then asks again. Models correct most mistakes when told precisely what was wrong.
The loop is bounded. After three attempts, it raises ExtractionError. Each retry is another paid call and more waiting, so do not raise the limit casually. If the first attempt fails often, fix the prompt or the schema instead.
The myth: Structured Outputs guarantee correct data
Since constrained generation arrived, a common claim is that the validation problem is solved. Half of it is. The model can no longer return a malformed object. It can still return a perfectly formed lie.
Consider the fields in the ticket. A schema can state that customer_email is a string or null. It cannot state that the address was actually in the message. A schema can state that priority is an integer. It cannot know that a double charge deserves a 4 and not a 1.
The original insight: check extracted values against the source
Here is the check I add to every extraction pipeline, and it costs one line. For any field that should be copied from the input, such as an email, an order number, a date, or a name, verify that the value appears in the source text.
check_grounding does that for the email. It is a substring test, it needs no model, and it catches the most damaging kind of error: an invented identifier that looks right. A hallucinated summary is embarrassing. A hallucinated account number sends a refund to the wrong customer.
Think of validation in three layers, and use all of them:
| Layer | Question | Tool |
|---|---|---|
| Shape | Are the fields and types right? | JSON Schema, constrained generation, Pydantic parsing |
| Rules | Are the values allowed? | Pydantic validators |
| Grounding | Do copied values exist in the input? | Plain Python checks against the source |
OpenAI Structured Outputs with a Pydantic model
The OpenAI SDK accepts a Pydantic class directly and returns a parsed instance. This call is billed. For client setup, see this guide to OpenAI client basics, timeouts, and retries.
from openai import OpenAI
from extract import OPENAI_MODEL, ExtractionError, Ticket, check_grounding
def extract_openai(client: OpenAI, text: str) -> Ticket:
completion = client.beta.chat.completions.parse(
model=OPENAI_MODEL,
messages=[
{'role': 'system', 'content': 'Extract a support ticket from the user message.'},
{'role': 'user', 'content': text},
],
response_format=Ticket,
temperature=0,
)
message = completion.choices[0].message
if message.refusal:
raise ExtractionError(f'model refused: {message.refusal}')
ticket = message.parsed
problems = check_grounding(ticket, text)
if problems:
raise ExtractionError('; '.join(problems))
return ticket
if __name__ == '__main__':
from extract import MESSAGE
print(extract_openai(OpenAI(timeout=30.0, max_retries=2), MESSAGE))
You no longer paste the schema into the prompt. The SDK converts the class to a strict JSON Schema and sends it as the response format. Three practical details matter.
- Refusals are separate. If the model declines for safety reasons,
parsedis empty andrefusalholds the explanation. Check it first. - Strict schemas have limits. Every field must be required, and extra properties must be forbidden. Constraint keywords such as minimum, maximum, and string length are not supported in strict mode at the time of writing. That is why the range rule lives in a Python validator, not in
Field(ge=1, le=5). - Truncation raises. If the reply hits the token limit, the JSON is incomplete, and the SDK raises an error instead of returning half an object.
Structured Outputs need a model that supports them. The snapshot named in the constant does. Keep the name in one place, because the list of supporting models will grow.
JSON mode is the older, weaker option: response_format={'type': 'json_object'}. It ensures valid JSON syntax and nothing about the fields. Prefer Structured Outputs where available, and keep the repair loop for JSON mode.
Anthropic: extraction through a forced tool call
Claude does not have a dedicated structured output mode today. The established pattern is tool calling: define one tool whose input schema is your model, and force the model to call it. This call is billed too.
from anthropic import Anthropic
from pydantic import ValidationError
from extract import ANTHROPIC_MODEL, ExtractionError, Ticket, check_grounding
def extract_anthropic(client: Anthropic, text: str) -> Ticket:
message = client.messages.create(
model=ANTHROPIC_MODEL,
max_tokens=500,
temperature=0,
tools=[
{
'name': 'record_ticket',
'description': 'Record the support ticket extracted from the message.',
'input_schema': Ticket.model_json_schema(),
}
],
tool_choice={'type': 'tool', 'name': 'record_ticket'},
messages=[{'role': 'user', 'content': text}],
)
block = next(item for item in message.content if item.type == 'tool_use')
try:
ticket = Ticket.model_validate(block.input)
except ValidationError as error:
raise ExtractionError(str(error)) from error
problems = check_grounding(ticket, text)
if problems:
raise ExtractionError('; '.join(problems))
return ticket
Nothing executes the “tool”. It exists only to give the model a schema to fill. tool_choice forces the call, so the reply always contains a tool_use block. The arguments usually match the schema, yet this is not a hard guarantee, so the Pydantic validation step is not optional here. For frequent failures, wrap the call in the same repair idea as the offline example.
Local and smaller models follow the same logic. Where no constrained mode exists, put the schema in the prompt, parse with Pydantic, and repair.
Test without the network
extract_with_repair takes the generator as an argument, so tests pass plain functions. Save this as test_extract.py.
import json
import pytest
from extract import MESSAGE, ExtractionError, extract_with_repair
GOOD = {
'category': 'billing',
'priority': 4,
'summary': 'Customer was charged twice.',
'customer_email': 'dana@example.com',
'needs_human': True,
}
def replay(*replies: str):
queue = list(replies)
seen: list[list[dict]] = []
def generate(messages: list[dict]) -> str:
seen.append(list(messages))
return queue.pop(0)
generate.seen = seen
return generate
def test_valid_reply_is_returned_on_the_first_attempt():
generate = replay(json.dumps(GOOD))
ticket = extract_with_repair(generate, MESSAGE)
assert ticket.priority == 4
assert len(generate.seen) == 1
def test_prose_around_json_triggers_a_repair():
generate = replay('Sure! Here is the JSON: {}', json.dumps(GOOD))
assert extract_with_repair(generate, MESSAGE).category == 'billing'
assert 'rejected' in generate.seen[1][-1]['content']
def test_invented_email_is_rejected():
bad = json.dumps(GOOD | {'customer_email': 'someone.else@example.com'})
generate = replay(bad, json.dumps(GOOD))
assert extract_with_repair(generate, MESSAGE).customer_email == 'dana@example.com'
def test_gives_up_after_the_attempt_limit():
bad = json.dumps(GOOD | {'priority': 9})
with pytest.raises(ExtractionError):
extract_with_repair(replay(bad, bad), MESSAGE, max_attempts=2)
python -m pytest -q test_extract.py
.... [100%]
4 passed in 0.14s
These tests cover the paths that matter: success, malformed text, an invented value, and exhaustion. They run in milliseconds and cost nothing. To judge how well a real model fills the fields, keep a separate set of real messages with expected tickets, and run it on purpose against the live API.
Design schemas that models fill well
- Keep them shallow and small. Five to ten fields at one level work far better than deep nesting. Split a large extraction into two calls.
- Prefer enumerations to free text. A
Literalwith four values is checkable. A free-text “type” field is not. - Name fields plainly.
customer_emailtells the model what to put there.field_3does not. - Describe units and formats. “Amount in cents as an integer” removes a whole class of errors.
- Order fields deliberately. The model writes fields in schema order. A reasoning or evidence field placed before the conclusion gives the model room to think before it commits.
- Allow null. If information may be absent, say so in the type. Otherwise the model invents something to fill the slot.
Cost and latency
The schema is sent with every request, so a large schema adds input tokens to every call. With Structured Outputs, the first request for a new schema is slower while the provider prepares it, and later requests are fast. Each repair attempt is another full call. Log the number of attempts per request. A rising repair rate is your earliest sign that a prompt, a schema, or a model version has drifted.
How real systems use structured output
- Validation at the boundary. Model output is parsed into a Pydantic object in one function. Nothing downstream ever sees raw model text.
- Deterministic follow-up. The model fills fields, and ordinary code acts on them. A router sends tickets by
categorywith anifstatement, not with another prompt. - Confidence by rule, not by asking. Systems flag records for human review when grounding checks fail or required fields are null. They do not rely on the model’s own confidence score.
- Typed APIs. A FastAPI endpoint can use the same Pydantic class as its response model, so the extraction result flows to clients without conversion.
- Stored raw replies. Teams keep the original model text beside the parsed object, with access controls, so that a wrong record can be investigated later.
We once hit a bug when an invoice extraction service returned a well-formed record with an invoice number that did not exist in the document. The schema was satisfied, and every type was right. The model had filled a required field from a similar-looking reference elsewhere on the page. A downstream system matched payments on that number, so three payments were attached to the wrong invoice before anyone noticed. A substring check against the source text would have rejected the record instantly, and we added one that day.
Choosing an approach: a decision framework
- Does your provider offer schema-constrained generation? Use it. It removes shape errors entirely.
- Does it offer tool calling but no constrained mode? Use a forced tool call, and validate the arguments.
- Neither? Put the schema in the prompt, parse with Pydantic, and use a bounded repair loop.
- Are any fields copied from the input? Add grounding checks for them, in every case above.
- Do values have business limits? Enforce them with Pydantic validators, since strict schemas may not accept those keywords.
- Is a wrong record costly? Route failures and nulls to a person, and never fill gaps with defaults.
When NOT to use structured output
- The answer is prose for a person. A reply to a customer or a summary to read does not need a schema. Forcing one makes the writing stiffer.
- The input is already structured. Parsing a CSV file or a known log format is a job for a parser. It is faster, free, and exact.
- The extraction needs perfect accuracy with no review. For amounts in contracts or medication doses, model extraction is a draft for a human to confirm, however clean the JSON looks.
Common mistakes
- Parsing with json.loads and trusting the result. Valid JSON can have missing fields and wrong types. Errors surface far from the cause, as a
KeyErrorin unrelated code. - Treating a valid schema as a correct answer. Invented values pass shape checks and flow into your database as facts.
- Unbounded retries. A reply that can never validate loops until a timeout, and bills you for each attempt.
- No escape value in an enumeration. Without
'other'or null, the model must mislabel anything that does not fit. - Deeply nested schemas. Accuracy drops as nesting grows, and repair messages become hard for the model to act on.
- Putting range constraints only in the schema. Strict modes may reject or ignore them. Enforce limits in Python as well.
Key takeaways
- Define the output once as a Pydantic model, and use it as schema, parser, and rule set.
- Prefer schema-constrained generation, then forced tool calls, then prompt plus repair.
- A schema guarantees shape. It never guarantees truth.
- Check copied values against the source text with plain Python.
- Feed validation errors back to the model, and cap the number of attempts.
- Keep schemas small and flat, with clear names, descriptions, and an escape value.
- Test the pipeline with scripted replies, and evaluate real model quality separately.
FAQ
How do I get JSON output from an LLM in Python?
Define a Pydantic model for the result, pass its JSON Schema to the model through the provider’s structured output or tool calling feature, and parse the reply with model_validate_json. Validate before using any value.
What is the difference between JSON mode and Structured Outputs?
JSON mode guarantees that the reply is syntactically valid JSON, with any fields. Structured Outputs constrain generation so that the reply matches your specific JSON Schema, including required fields and types.
Can I use Pydantic with the OpenAI API?
Yes. The OpenAI Python SDK accepts a Pydantic class as response_format in its parse helper and returns a parsed instance on message.parsed.
How do I get structured output from Claude?
Define a single tool whose input_schema is your JSON Schema, and force it with tool_choice. Read the tool_use block from the response, and validate its input with Pydantic.
Do Structured Outputs prevent hallucinations?
No. They ensure the output has the right shape. The values can still be wrong or invented, so you need validators and checks against the source text.
Constrain the shape, verify the substance
Structured output turns a language model into a component you can call from ordinary code. The schema handles form, and providers now enforce it for you. What remains is the part no schema can do: deciding whether the values make sense and whether they came from the input at all. Write those checks once, next to the model class.
Rule of thumb: let the schema reject what is malformed, and let your own code reject what is merely plausible.
