Pydantic v2 Tutorial: Data Validation and Settings in Python
Validate untrusted data with typed models, write field and model validators, serialize with model_dump, and load typed settings from environment variables.
A signup endpoint accepted a JSON body and stored it. One client sent "age": "twenty". The value travelled through three functions, reached a database column of type integer, and failed there with a driver error that mentioned neither the field nor the request. A validation layer at the entrance would have rejected the request in microseconds, with a message naming the field.
Pydantic is the most widely used data validation library for Python. This Pydantic v2 tutorial covers version 2, whose final release arrived on 30 June 2023. You declare the shape of your data as a class with type annotations, and Pydantic checks and converts incoming data to match.
You need Python 3.11 and Pydantic 2.1. You should be comfortable with Python type hints, because a Pydantic model is built from them. Unlike ordinary annotations, which Python ignores at runtime, Pydantic enforces them.
My position: use dataclasses for plain data inside your program, and Pydantic at the boundary where data enters. Validating the same trusted object again and again is wasted work that makes code slower and no safer.
Install Pydantic v2 and write a first model
python --version
python -m venv .venv
source .venv/bin/activate # macOS and Linux
.venv\Scripts\Activate.ps1 # Windows PowerShell
python -m pip install "pydantic>=2.1,<2.3" pydantic-settings
python -c "import pydantic; print(pydantic.VERSION)"
2.1.1
A model is a class that inherits from BaseModel. Each annotated attribute is a field.
from pydantic import BaseModel
class Task(BaseModel):
id: int
title: str
done: bool = False
task = Task.model_validate({'id': '7', 'title': 'Write docs'})
print(task)
print(task.id + 1)
id=7 title='Write docs' done=False
8
model_validate() takes a dictionary and returns a model instance, or raises ValidationError. Notice that the string '7' became the integer 7. Fields with a default are optional, and fields without one are required. You can also construct a model with keyword arguments, Task(id=7, title='Write docs'), which runs the same validation.
The myth: Pydantic checks types strictly
Many developers assume that count: int accepts only integers. By default, Pydantic runs in lax mode, where it converts reasonable inputs to the declared type.
from pydantic import BaseModel, ConfigDict, ValidationError
class Lax(BaseModel):
count: int
active: bool
class Strict(BaseModel):
model_config = ConfigDict(strict=True)
count: int
active: bool
data = {'count': '42', 'active': 'yes'}
print(Lax.model_validate(data))
try:
Strict.model_validate(data)
except ValidationError as error:
print(error)
count=42 active=True
2 validation errors for Strict
count
Input should be a valid integer [type=int_type, input_value='42', input_type=str]
For further information visit https://errors.pydantic.dev/2.1/v/int_type
active
Input should be a valid boolean [type=bool_type, input_value='yes', input_type=str]
For further information visit https://errors.pydantic.dev/2.1/v/bool_type
Lax mode exists for good reasons. Query strings, form fields, and environment variables are always text, so '42' must become 42. However, the same leniency can hide a client bug when the input is JSON, where a number and a string are different things.
Here is the heuristic I use: lax for text channels, strict for typed channels. Environment variables, query parameters, and CSV cells need conversion. JSON from another service does not, and strict=True there turns a silent conversion into a visible contract violation.
Version 2 is already stricter than version 1 in one notable way. It no longer converts numbers to strings. A field declared as str that receives 123 fails with Input should be a valid string.
Constrain fields and write validators
Types alone rarely capture the rules. A title should not be empty, and a priority should stay within a range. Pydantic offers three layers, from declarative to custom.
| Layer | Tool | Use it for |
|---|---|---|
| Constraints | Field(min_length=3, ge=1, le=5) |
Lengths, ranges, patterns |
| Field validator | @field_validator('tags') |
Custom checks or cleanup for one field |
| Model validator | @model_validator(mode='after') |
Rules that involve several fields |
The complete example below uses all three. Save it as models.py.
from datetime import date
from typing import Annotated
from pydantic import (
BaseModel,
ConfigDict,
Field,
ValidationError,
field_validator,
model_validator,
)
Priority = Annotated[int, Field(ge=1, le=5)]
class TaskCreate(BaseModel):
model_config = ConfigDict(extra='forbid', str_strip_whitespace=True)
title: str = Field(min_length=3, max_length=200)
priority: Priority = 3
due: date | None = None
tags: list[str] = []
done: bool = False
@field_validator('tags')
@classmethod
def normalize_tags(cls, tags: list[str]) -> list[str]:
return sorted({tag.strip().lower() for tag in tags if tag.strip()})
@model_validator(mode='after')
def urgent_tasks_need_a_due_date(self) -> 'TaskCreate':
if self.priority == 5 and self.due is None:
raise ValueError('priority 5 tasks need a due date')
return self
def main() -> None:
task = TaskCreate.model_validate(
{'title': ' Write docs ', 'priority': '4', 'tags': ['Docs', ' docs', 'Writing']}
)
print(task)
print(task.model_dump())
print(task.model_dump_json())
try:
TaskCreate.model_validate({'title': 'ab', 'priority': 9, 'owner': 'asha'})
except ValidationError as error:
for item in error.errors():
print(item['loc'], item['type'], item['msg'])
if __name__ == '__main__':
main()
python models.py
title='Write docs' priority=4 due=None tags=['docs', 'writing'] done=False
{'title': 'Write docs', 'priority': 4, 'due': None, 'tags': ['docs', 'writing'], 'done': False}
{"title":"Write docs","priority":4,"due":null,"tags":["docs","writing"],"done":false}
('title',) string_too_short String should have at least 3 characters
('priority',) less_than_equal Input should be less than or equal to 5
('owner',) extra_forbidden Extra inputs are not permitted
Several things happened in the valid case. str_strip_whitespace trimmed the title. The string '4' became an integer. The field validator lowercased the tags, removed the duplicate, and sorted them.
In the invalid case, Pydantic reported every problem at once, not only the first. Each error has a location, a stable machine-readable type, and a human message. That list is exactly what an API should return to a client. The unknown key owner was rejected because of extra='forbid'. Without that setting, Pydantic ignores unknown keys silently, which hides typos such as priorty.
Details that trip people up
- Validators need
@classmethod. A field validator runs before the instance exists, so it receives the class and the value. Place@classmethodbelow@field_validator. - Raise
ValueError, notValidationError. Pydantic catchesValueErrorandAssertionErrorfrom your validator and turns them into a proper error entry. - Before and after modes. A field validator runs after type conversion by default, so
normalize_tagsalready receives a list of strings. Usemode='before'to see the raw input. - Model validators run last. An
aftermodel validator runs only when every field passed, so it can trustself. - Mutable defaults are safe. Pydantic copies the default
[]for each instance, unlike a plain function argument.
Required, optional, and nullable are three different things
Version 2 changed this rule, and it catches almost everyone who migrates.
| Declaration | Must be provided? | May be None? |
|---|---|---|
note: str |
Yes | No |
note: str = 'n/a' |
No | No |
note: str | None |
Yes | Yes |
note: str | None = None |
No | Yes |
In version 1, Optional[str] without a default quietly meant “optional, defaulting to None“. In version 2, it means required but nullable, and omitting it produces Field required. A field is optional only if it has a default.
Serialize with model_dump
Validation is half the job. Sending data out again is the other half.
task.model_dump() # dict of Python objects
task.model_dump(mode='json') # dict of JSON-safe values (dates become strings)
task.model_dump_json() # JSON string
task.model_dump(exclude={'done'}) # leave fields out
task.model_dump(exclude_unset=True) # only fields the caller actually sent
exclude_unset=True is the key to partial updates. It tells you which fields a client provided, so that you can change only those and leave the rest alone. To parse JSON text directly, use TaskCreate.model_validate_json(raw), which is faster than calling json.loads first.
Build models from objects
APIs often convert a database object into a response model. Enable from_attributes, and model_validate reads attributes instead of dictionary keys. The class below stands in for an ORM row, such as those produced by SQLAlchemy 2.0 ORM models.
from pydantic import BaseModel, ConfigDict
class TaskRow:
def __init__(self) -> None:
self.id = 1
self.title = 'Write docs'
self.done = False
self.internal_notes = 'do not expose'
class TaskRead(BaseModel):
model_config = ConfigDict(from_attributes=True)
id: int
title: str
done: bool
print(TaskRead.model_validate(TaskRow()))
id=1 title='Write docs' done=False
The response model lists only the fields you intend to publish, so internal_notes never leaves the server. Separate input models (TaskCreate) from output models (TaskRead). They look repetitive at first, and they prevent both over-posting and data leaks.
Load typed settings from the environment
Configuration is untrusted input too. It arrives as text from environment variables, and a typo in one can break production. In version 2, settings support lives in a separate package, pydantic-settings. Save this as settings.py.
from pydantic import Field, SecretStr
from pydantic_settings import BaseSettings, SettingsConfigDict
class Settings(BaseSettings):
model_config = SettingsConfigDict(env_prefix='TASKAPI_', env_file='.env')
database_url: str = 'sqlite:///tasks.db'
debug: bool = False
page_size: int = Field(default=20, ge=1, le=100)
api_key: SecretStr
def main() -> None:
settings = Settings()
print(settings)
if __name__ == '__main__':
main()
Each field is read from an environment variable with the prefix: TASKAPI_DEBUG, TASKAPI_PAGE_SIZE, and so on. Run it without the required key first:
python settings.py
pydantic_core._pydantic_core.ValidationError: 1 validation error for Settings
api_key
Field required [type=missing, input_value={}, input_type=dict]
The program refuses to start, and the message names the missing setting. Now provide values:
# macOS and Linux
TASKAPI_API_KEY=sk-your-key-here TASKAPI_DEBUG=true python settings.py
# Windows PowerShell
$env:TASKAPI_API_KEY = 'sk-your-key-here'; $env:TASKAPI_DEBUG = 'true'; python settings.py
database_url='sqlite:///tasks.db' debug=True page_size=20 api_key=SecretStr('**********')
SecretStr masks the value whenever the object is printed or logged. To use the real value, call settings.api_key.get_secret_value(). Real environment variables take priority over the .env file, and you should keep that file out of version control.
Migrating from Pydantic v1
Most of the work in a migration is renaming. The old names still exist in 2.x and emit deprecation warnings.
| Pydantic v1 | Pydantic v2 |
|---|---|
Model.parse_obj(data) |
Model.model_validate(data) |
Model.parse_raw(text) |
Model.model_validate_json(text) |
obj.dict() |
obj.model_dump() |
obj.json() |
obj.model_dump_json() |
@validator |
@field_validator |
@root_validator |
@model_validator |
class Config: |
model_config = ConfigDict(...) |
orm_mode = True |
from_attributes=True |
from pydantic import BaseSettings |
from pydantic_settings import BaseSettings |
The Pydantic team publishes a tool, bump-pydantic, that applies many of these renames automatically. Behavior changes need human review: the optional-field rule above, the end of number-to-string conversion, and validator signatures. The project says the new Rust core makes validation between 5 and 50 times faster than version 1. Treat that as the maintainers’ claim, and measure your own models if speed is your reason to migrate.
Pydantic models versus dataclasses
| Question | Dataclass | Pydantic model |
|---|---|---|
| Validates at runtime | No | Yes |
| Converts input types | No | Yes, in lax mode |
| Construction cost | Very low | Higher, because validation runs |
| JSON in and out | Manual | Built in |
| Dependency | Standard library | Third-party package |
| Best for | Trusted data inside your program | Untrusted data at the boundary |
For a deeper look at the left column, see this guide to dataclasses. The two tools complement each other. Data enters through a Pydantic model, and after that point your code can rely on the types.
How real systems use Pydantic
- Request and response schemas. Web frameworks such as FastAPI validate each request body against a model and serialize responses through another. Invalid requests receive a 422 response built from
error.errors(). - Settings loaded once at startup. The application builds a
Settingsobject when it boots. A missing or malformed variable stops the deploy, instead of failing on the first request that needs it. - Messages from queues and webhooks. Consumers validate each payload on arrival. Bad messages go to a dead-letter queue with the error list attached.
- Separate models per direction. Teams keep
Create,Update, andReadmodels, so that clients cannot set server-owned fields and responses expose only intended data. - Plain objects in the core. Business logic receives validated values and does not re-validate them on every call.
We once hit a bug when a service started with TASKAPI_PAGE_SIZE=2O, with a capital letter O instead of a zero, in one environment. The old code read the variable with os.environ.get and called int() lazily, so the service booted cleanly and crashed only on list endpoints. After moving to a settings model, the same typo stopped the deploy at startup with an error naming page_size. A failure at boot is far cheaper than one found by users.
Choosing where and how to validate: a decision framework
- Does this data come from outside the process? Requests, files, queues, environment variables, and third-party APIs get a Pydantic model.
- Is it created by your own code from already validated values? Use a dataclass or a plain object. Do not validate it again.
- Is the input channel text-only? Keep lax mode, so that
'42'converts. - Is it JSON from a system that should send correct types? Use
strict=Trueto catch contract drift. - Should unknown fields be an error? For input models, set
extra='forbid'. For models that read third-party responses, leave the default, so new fields do not break you.
When NOT to use Pydantic
- Internal data structures in hot paths. Creating millions of validated objects in a loop pays the validation cost millions of times. Use dataclasses or tuples for trusted data.
- Large tabular data. Validating a million-row CSV one model at a time is slow. Load it with a dataframe library and validate columns in bulk.
- Small scripts with no outside input. A script that reads constants defined in the same file gains a dependency and nothing else.
Common mistakes
- Assuming Optional means optional. In version 2,
str | Nonewithout a default is required. Requests that omitted the field start failing withField required. - Forgetting @classmethod on a field validator. Pydantic raises an error when the class is defined, and the message is not obvious the first time.
- Using one model for input and output. Clients can then set fields such as
idoris_admin, and responses can leak internal fields. - Ignoring extra fields on input. With the default setting, a misspelled key is dropped without a word. The client believes it set a value that never arrived.
- Printing or logging secrets as plain strings. A
strfield for an API key appears in every log line that prints the settings. UseSecretStr. - Calling old v1 methods.
.dict()and.parse_obj()still work in 2.x but warn, and they are scheduled for removal in a later major version.
Key takeaways
- Define the shape of incoming data as a
BaseModel, and parse withmodel_validate(). - Pydantic converts types by default. Use
strict=Truewhere conversion would hide a bug. - Layer your rules:
Fieldconstraints, then field validators, then model validators. - A field is optional only when it has a default.
- Serialize with
model_dump()andmodel_dump_json(), and use separate models for input and output. - Load configuration through
pydantic-settings, so that bad settings fail at startup. - Validate at the boundary once, and use plain typed objects inside.
FAQ
What is Pydantic used for?
Pydantic validates and converts data using Python type annotations. Developers use it to check API requests, parse JSON, and load configuration into typed objects.
What changed in Pydantic v2?
Version 2 moved the validation core to Rust for speed, renamed methods to names such as model_validate and model_dump, replaced @validator with @field_validator, and moved BaseSettings to the pydantic-settings package.
What is the difference between model_validate and model_dump?
model_validate() turns input data into a validated model instance. model_dump() does the reverse and turns a model instance into a dictionary.
Should I use Pydantic or dataclasses?
Use Pydantic for data that enters your program from outside, because it validates at runtime. Use dataclasses for trusted data inside your program, where validation would only add cost.
How do I read environment variables with Pydantic v2?
Install pydantic-settings, define a class that inherits from BaseSettings, and create an instance. Each field is filled from the matching environment variable and validated against its type.
Validate at the boundary, trust inside
Pydantic gives you one place to state what valid data looks like, and it enforces that statement at runtime. Put that place where data enters: requests, files, messages, and settings. Everything behind it can then be simpler, because it no longer has to doubt its inputs.
Rule of thumb: parse untrusted data exactly once, as early as possible, and never pass a raw dictionary deeper than the function that received it.
