Python Dataclasses and Classes: When to Use Each
Dataclasses remove boilerplate for plain data, and regular classes protect state and behavior. Learn frozen, slots, field(), and a rule for choosing.
You write a class with five attributes. You add an __init__ that copies five parameters, a __repr__ that formats five values, and an __eq__ that compares five pairs. That is about twenty lines that say the same thing three times. Six months later someone adds a sixth attribute and forgets to update __eq__, and two different objects start comparing as equal.
Python dataclasses solve this. The @dataclass decorator, from the standard library’s dataclasses module, reads the annotated fields of a class and writes those methods for you. The class remains an ordinary Python class in every other way.
This guide uses Python 3.10, which added the slots and kw_only options. Check your interpreter with python --version. You should already know how to define a function. A class, in one sentence, is a template that bundles data (attributes) with functions that work on it (methods).
My position: reach for a dataclass first whenever an object mostly carries values, and write a regular class only when you need to guard an invariant or hide state. Most codebases get this backwards and hand-write boilerplate they never needed.
What a regular class makes you write
Start with the manual version, so you can see what the decorator replaces.
class Point:
def __init__(self, x: float, y: float) -> None:
self.x = x
self.y = y
a = Point(1.0, 2.0)
b = Point(1.0, 2.0)
print(a)
print(a == b)
<__main__.Point object at 0x7f8e5c2a3d90>
False
The output is unhelpful in a log, and two points with equal coordinates are not equal. By default, == on class instances compares identity: whether both names refer to the same object. To fix both problems by hand, you must write __repr__ and __eq__ yourself.
How Python dataclasses remove the boilerplate
The dataclass version declares the fields once. The annotations are required, because the decorator finds fields by looking for them.
from dataclasses import dataclass
@dataclass
class Point:
x: float
y: float
a = Point(1.0, 2.0)
b = Point(1.0, 2.0)
print(a)
print(a == b)
Point(x=1.0, y=2.0)
True
The decorator generated three methods: an __init__ that accepts x and y, a readable __repr__, and an __eq__ that compares field by field. The work happens once, when Python defines the class, so instances are as fast as hand-written ones.
Defaults and default_factory
Give a field a default by assigning a value. For a mutable default such as a list, you must use field(default_factory=...). A dataclass refuses the unsafe form outright:
# WRONG: raises at class definition time
@dataclass
class Article:
title: str
tags: list[str] = []
ValueError: mutable default <class 'list'> for field tags is not allowed: use default_factory
# RIGHT: each instance gets its own new list
from dataclasses import dataclass, field
@dataclass
class Article:
title: str
tags: list[str] = field(default_factory=list)
This protection exists because one shared list would leak data between instances. It is the same problem as the mutable default argument trap in functions, except that here Python stops you early.
Fields without defaults must come before fields with defaults. Otherwise you get TypeError: non-default argument 'title' follows default argument.
Validate in __post_init__
The generated __init__ only assigns values. To check them, define __post_init__, which runs right after the assignments.
from dataclasses import dataclass
@dataclass
class Percentage:
value: float
def __post_init__(self) -> None:
if not 0 <= self.value <= 100:
raise ValueError(f'value must be between 0 and 100, got {self.value}')
However, this check runs only at construction. Unless the dataclass is frozen, later code can still assign p.value = 500 and skip it.
The options that matter: frozen, slots, kw_only, order
The decorator accepts keyword options. Four of them cover nearly every practical need.
| Option | Effect | Cost | Available since |
|---|---|---|---|
frozen=True |
Blocks attribute assignment after creation, and makes instances hashable | You must create a new object to change a value | 3.7 |
slots=True |
Stores fields in fixed slots: less memory, and typos in attribute names fail | No ad hoc attributes, and some multiple-inheritance limits | 3.10 |
kw_only=True |
Callers must pass every field by name | More verbose construction | 3.10 |
order=True |
Adds <, >, and sorting, comparing fields in order |
Ordering by field position may not match your domain | 3.7 |
Each option turns a silent mistake into a loud one. A frozen instance rejects assignment:
dataclasses.FrozenInstanceError: cannot assign to field 'x'
A slotted instance rejects a misspelled attribute, which a normal class would accept silently as a new attribute:
AttributeError: 'Point' object has no attribute 'z'
The hash rule nobody expects
Here is a connection the documentation states but most people miss until it bites. A plain @dataclass defines __eq__, and Python then sets __hash__ to None. The instance becomes unhashable, so you cannot put it in a set or use it as a dictionary key:
@dataclass
class Point:
x: float
y: float
seen = {Point(1.0, 2.0)}
TypeError: unhashable type: 'Point'
The hand-written class at the top of this article was hashable, so converting it to a dataclass can break code that stored instances in a set. The fix is frozen=True. A frozen dataclass gets a hash computed from its fields, which is safe because those fields cannot change.
My heuristic follows from this: default to @dataclass(frozen=True, slots=True) and remove an option only when you have a reason. Immutable, slotted records are hashable, smaller, and safe to share between threads and caches.
Measure what slots saves
This is my own illustrative test. Save it and run it to see the numbers on your machine.
import tracemalloc
from dataclasses import dataclass
@dataclass
class Plain:
x: int
y: int
@dataclass(slots=True)
class Slotted:
x: int
y: int
def measure(cls) -> float:
tracemalloc.start()
items = [cls(i, i) for i in range(100_000)]
current, _peak = tracemalloc.get_traced_memory()
tracemalloc.stop()
del items
return current / 1_000_000
if __name__ == '__main__':
print(f'plain: {measure(Plain):.1f} MB')
print(f'slotted: {measure(Slotted):.1f} MB')
On my machine, the slotted version uses well under half the memory of the plain one. A normal instance carries a per-object dictionary for its attributes, and slots remove it. The saving matters when you hold hundreds of thousands of objects, and it is irrelevant for a handful.
The myth: dataclass type hints validate your data
Many developers assume that x: float in a dataclass means Python checks that x is a float. It does not. The annotation tells the decorator that a field exists, and it tells a type checker what you intend. At runtime, nothing enforces it.
from dataclasses import dataclass
@dataclass
class Point:
x: float
y: float
p = Point('left', None)
print(p)
Point(x='left', y=None)
No error appears. Consequently, a dataclass is a good fit for data your own code creates and a poor fit for unchecked input such as JSON from a request. For that boundary, validate explicitly in __post_init__, or use a validation library such as pydantic or attrs with validators.
When a regular class is the right tool
A regular class earns its extra lines when the object’s job is to protect state, not to expose it. Consider a bank account. The balance must never go negative, and callers must not set it directly.
class Account:
def __init__(self, owner: str, opening_balance: int = 0) -> None:
if opening_balance < 0:
raise ValueError('opening balance cannot be negative')
self.owner = owner
self._balance = opening_balance
@property
def balance(self) -> int:
return self._balance
def withdraw(self, amount: int) -> None:
if amount <= 0:
raise ValueError('amount must be positive')
if amount > self._balance:
raise ValueError('insufficient funds')
self._balance -= amount
def __repr__(self) -> str:
return f'Account(owner={self.owner!r}, balance={self._balance})'
The leading underscore in _balance marks the attribute as internal by convention. The @property exposes a read-only view. Every change goes through a method that enforces the rule. A dataclass would generate an __init__ and an __eq__ that expose and compare exactly the state you want to hide.
Dataclasses compared with the alternatives
| Option | Named fields | Mutable | Methods | Best for |
|---|---|---|---|---|
| Tuple | No | No | No | Two or three values returned from a function |
dict |
String keys, unchecked | Yes | No | Dynamic keys and JSON-shaped data |
typing.NamedTuple |
Yes | No | Yes | Small records that must also unpack like tuples |
@dataclass |
Yes | Your choice | Yes | Plain data with a known set of fields |
| Regular class | Yes | Yes | Yes | Behavior, invariants, and hidden state |
attrs (third party) |
Yes | Your choice | Yes | Dataclass features plus validators and converters |
If plain tuples and dictionaries already serve you, keep them for small local data. Move to a dataclass when the same shape crosses function boundaries and you find yourself remembering what position 3 means.
A complete example: both in one program
Real programs use both tools together. In this small task tracker, Task is plain data, so it is a frozen dataclass. TaskBoard owns the collection and the ID counter, so it is a regular class. Save the file as board.py.
from dataclasses import asdict, dataclass, replace
from datetime import date
@dataclass(frozen=True, slots=True, kw_only=True)
class Task:
id: int
title: str
due: date | None = None
tags: tuple[str, ...] = ()
done: bool = False
def __post_init__(self) -> None:
if not self.title.strip():
raise ValueError('title must not be empty')
class TaskBoard:
def __init__(self) -> None:
self._tasks: dict[int, Task] = {}
self._next_id = 1
def add(self, title: str, *, due: date | None = None, tags=()) -> Task:
task = Task(id=self._next_id, title=title, due=due, tags=tuple(tags))
self._tasks[task.id] = task
self._next_id += 1
return task
def complete(self, task_id: int) -> Task:
task = replace(self._tasks[task_id], done=True)
self._tasks[task_id] = task
return task
@property
def open_count(self) -> int:
return sum(1 for task in self._tasks.values() if not task.done)
def main() -> None:
board = TaskBoard()
first = board.add('Write release notes', tags=['docs'])
board.add('Rotate API keys', tags=['security'])
print(first)
print(board.complete(first.id))
print(board.open_count)
print(asdict(first))
if __name__ == '__main__':
main()
python board.py
Task(id=1, title='Write release notes', due=None, tags=('docs',), done=False)
Task(id=1, title='Write release notes', due=None, tags=('docs',), done=True)
1
{'id': 1, 'title': 'Write release notes', 'due': None, 'tags': ('docs',), 'done': False}
Three details are worth noting. First, replace() builds a modified copy, which is how you “change” a frozen object. Second, the original first still shows done=False, because nothing mutated it. Third, open_count counts with a generator expression, so it builds no temporary list. The tags field is a tuple, not a list, so the frozen task stays fully immutable and hashable.
How real systems structure classes and dataclasses
- Dataclasses as messages between layers. A function that parses a file returns dataclass instances, and the next layer consumes them. The fields are the contract between the two.
- Frozen dataclasses for configuration. Code builds settings once at startup and passes the object around. Nothing can change a timeout halfway through a request.
- Frozen dataclasses as cache and dictionary keys. A query described by
(user_id, date_range, filters)becomes a hashable value object. - Regular classes for services and resources. Database clients, repositories, and connection pools hold private state and expose methods. Equality between two of them is rarely meaningful.
- Slots for high-volume objects. Parsers and simulations that create millions of small records add
slots=Trueto cut memory.
A mistake I have seen in production is a shared settings object that a request handler modified “temporarily” to raise a retry limit. The change stayed in place for every later request in that worker process. Each worker behaved differently depending on its history, which made the bug nearly impossible to reproduce. Making the settings dataclass frozen turned that line into an immediate FrozenInstanceError in tests.
Choosing between a dataclass and a class: a decision framework
Answer these questions in order and stop at the first clear match.
- Must the object enforce a rule on every change? If state must stay consistent across method calls, write a regular class with private attributes.
- Does it wrap a resource such as a connection or file? Write a regular class. Generated equality and
reprmake no sense for it. - Is it a value defined entirely by its fields? Use
@dataclass(frozen=True, slots=True). Two instances with equal fields are interchangeable. - Is it a record that code fills in step by step? Use a mutable
@dataclass, and accept that it is unhashable. - Does the data arrive from outside your program? Parse and validate it first, then build the dataclass from checked values.
When NOT to use a dataclass
- When you need runtime validation of untrusted input. A dataclass stores whatever it receives. Request bodies and file contents need explicit checks or a validation library.
- When the keys are not known in advance. Headers, feature flags, and arbitrary JSON have dynamic keys. A dictionary models them honestly.
- When equality by fields is wrong. Two user sessions with identical field values are still different sessions. Generated
__eq__would merge them in a set or a comparison.
Common mistakes
- Forgetting the type annotation. A line such as
retries = 3without an annotation is a class attribute, not a field. It is missing from__init__,__repr__, and__eq__. - Using a mutable dataclass as a dictionary key. The default dataclass is unhashable and raises
TypeError. Addfrozen=True. - Trusting annotations to validate. Wrong types pass silently at construction and fail much later, far from the real cause.
- Putting a list inside a frozen dataclass. Freezing blocks reassignment, but code can still call
task.tags.append(). The object is also unhashable, because lists are. Use a tuple. - Turning on order=True without thinking. Comparison follows field order, so sorting tasks compares IDs first, then titles. Pass an explicit
keytosorted()instead. - Writing a regular class out of habit. Hand-written
__eq__and__repr__drift out of date when someone adds a field. Bugs appear as wrong comparisons.
Key takeaways
- Use a dataclass for objects that mostly hold values, and a regular class to guard state and behavior.
- Default to
@dataclass(frozen=True, slots=True), and relax it only for a reason. - Use
field(default_factory=list)for mutable defaults. - A plain dataclass is unhashable, and freezing it makes it hashable.
- Annotations declare fields. They do not validate anything at runtime.
- Change a frozen instance with
dataclasses.replace(), which returns a copy. - Add
kw_only=Truewhen a class has several fields of the same type.
FAQ
What is a dataclass in Python?
A dataclass is a regular class decorated with @dataclass. The decorator reads the annotated fields and generates __init__, __repr__, and __eq__ for you.
What is the difference between a dataclass and a regular class?
A dataclass writes the standard methods from your field list, so it suits plain data. A regular class leaves everything to you, which suits objects that hide state and enforce rules through methods.
Do Python dataclasses validate types?
No. Type hints in a dataclass are not checked at runtime. Add checks in __post_init__ or use a validation library if you need enforcement.
What does frozen=True do in a dataclass?
It makes instances immutable. Assigning to a field raises FrozenInstanceError, and the instance becomes hashable, so you can use it in sets and as a dictionary key.
Should I use a dataclass or a NamedTuple?
Use a dataclass in most cases. Choose NamedTuple only when the object must also behave like a tuple, for example when callers unpack it by position.
Data gets a dataclass, behavior gets a class
The choice is rarely close once you ask what the object is for. If it carries values from one place to another, declare the fields and let Python write the rest. If it protects something, write the class by hand and keep its state private.
Rule of thumb: if two objects with the same fields are the same thing, use a frozen dataclass; if they are not, write a class.
