Python Linting and Formatting: Ruff, Black, and flake8 Compared
Linters find bugs and formatters end style debates. Compare Ruff, Black, flake8, isort, and pylint, and copy a working pyproject.toml and pre-commit setup.
A pull request with a two-line bug fix arrives with 400 changed lines. The author’s editor re-wrapped every long line and re-ordered the imports. The reviewer cannot find the fix, approves anyway, and the diff hides an unrelated mistake. Automated formatting would have made that diff two lines long.
Python linting and formatting tools read your source code without running it. A linter reports problems, such as an unused import or a mutable default argument. A formatter changes whitespace, quotes, and line breaks to one consistent style. The two overlap a little and replace each other not at all.
This guide uses Python 3.11 with Ruff 0.0.26x, Black 23.x, and flake8 6.x. You need a terminal and a project with a virtual environment. Run python --version to confirm your interpreter.
My position: automate style completely and stop reviewing it by hand. Then spend your linting attention on the rules that find bugs, not on the ones that police taste.
Python linting and formatting: what each tool does
Five tools cover almost every Python project.
- Black is a formatter. It rewrites files to one style with very few options. Its default line length is 88 characters.
- isort sorts and groups
importstatements. - flake8 is a linter that combines pyflakes (logic errors), pycodestyle (PEP 8 style), and mccabe (complexity). Plugins add more rules.
- pylint is a deeper, slower linter. It infers types to find errors that simpler tools miss, and it has many opinionated checks.
- Ruff is a linter written in Rust. It reimplements the rules of flake8, many flake8 plugins, isort, and pyupgrade in one binary, and it can fix many violations automatically.
A type checker is a third category. It verifies that types line up across function calls, which no linter here does. See this guide to static type checking with mypy.
Set up and try the tools on one file
Install the tools into a virtual environment.
python --version
python -m venv .venv
source .venv/bin/activate # macOS and Linux
.venv\Scripts\Activate.ps1 # Windows PowerShell
python -m pip install ruff black flake8
Save this deliberately messy file as messy.py. It runs, and it contains six problems.
import os, sys
import json
from typing import List
def load(path, cache={}):
if path in cache: return cache[path]
try:
data = json.load(open(path))
except:
data = None
cache[path] = data
unused = 42
return data
def total(items: List[int]):
return sum([i for i in items])
What flake8 reports
python -m flake8 messy.py
messy.py:1:1: F401 'os' imported but unused
messy.py:1:1: F401 'sys' imported but unused
messy.py:1:10: E401 multiple imports on one line
messy.py:5:1: E302 expected 2 blank lines, found 1
messy.py:6:19: E701 multiple statements on one line (colon)
messy.py:9:5: E722 do not use bare 'except'
messy.py:12:5: F841 local variable 'unused' is assigned to but never used
messy.py:15:1: E302 expected 2 blank lines, found 1
The letter prefix tells you the source. F codes come from pyflakes and usually indicate real mistakes. E and W codes come from pycodestyle and concern style. Notice what is missing: flake8 says nothing about the mutable default cache={}. That check lives in a plugin, flake8-bugbear.
What Ruff reports
Ruff enables a small default rule set. Select more rule families on the command line to match and exceed the flake8 run:
ruff check --select E,F,B,UP messy.py
Trimmed output:
messy.py:1:1: E401 Multiple imports on one line
messy.py:1:8: F401 [*] `os` imported but unused
messy.py:1:12: F401 [*] `sys` imported but unused
messy.py:5:22: B006 Do not use mutable data structures for argument defaults
messy.py:6:20: E701 Multiple statements on one line (colon)
messy.py:9:5: E722 Do not use bare `except`
messy.py:12:5: F841 Local variable `unused` is assigned to but never used
messy.py:15:18: UP006 [*] Use `list` instead of `List` for type annotation
Ruff found the mutable default (B006, from its built-in bugbear rules) and suggested the modern annotation (UP006, from pyupgrade). The [*] marker means Ruff can fix the issue itself. Run ruff check --fix to apply those fixes.
What Black changes
python -m black messy.py
reformatted messy.py
All done! ✨ 🍰 ✨
1 file reformatted.
Black fixes layout only. It adds the blank lines, splits if path in cache: return cache[path] across two lines, and normalizes quotes. It does not split import os, sys, remove unused imports, or touch the bare except. Those belong to the linter. Running both tools on the same file shows the division of labor clearly.
Ruff, Black, flake8, isort, and pylint compared
| Tool | Job | Speed | Auto-fix | Config in pyproject.toml | Maturity |
|---|---|---|---|---|---|
| Black | Formatter | Fast | Yes, that is its job | Yes | Stable |
| isort | Import sorter | Fast | Yes | Yes | Stable |
| flake8 | Linter | Moderate | No | No. Uses .flake8, setup.cfg, or tox.ini |
Stable, large plugin ecosystem |
| pylint | Deep linter | Slow | No | Yes | Stable |
| Ruff | Linter and import sorter | Very fast | Yes, for many rules | Yes | Young, 0.0.x releases |
On speed, the Ruff project describes itself as 10 to 100 times faster than existing linters. Do not take a README’s word for it, or mine. Time both tools on your own repository:
time python -m flake8 src/
time ruff check src/
On a large codebase, the practical difference is that flake8 takes long enough to tempt people to skip it, while Ruff finishes before you notice it started. Speed matters because it changes behavior: a check that takes under a second can run on every save and every commit.
The myth: Ruff replaces Black
Because Ruff replaces flake8, isort, and several other tools, many people assume it replaces Black too. Today it does not. Ruff is a linter. It can fix individual violations, such as removing an unused import, but it does not reformat a file’s layout. You still need a formatter, and Black is the standard choice.
The reverse myth also circulates: that a formatter makes code “clean”. Black will format a function with a bare except and a mutable default argument beautifully. Formatting is about consistency, and it tells you nothing about correctness.
A configuration you can copy
Put tool settings in pyproject.toml, the standard configuration file at the project root.
[tool.black]
line-length = 88
target-version = ["py311"]
[tool.ruff]
line-length = 88
target-version = "py311"
select = ["E", "F", "I", "B", "UP"]
ignore = ["E501"]
[tool.ruff.per-file-ignores]
"tests/*" = ["B011"]
The select list turns on rule families:
| Prefix | Origin | What it finds |
|---|---|---|
F |
pyflakes | Unused imports and variables, undefined names |
E |
pycodestyle | PEP 8 style errors |
I |
isort | Unsorted imports |
B |
flake8-bugbear | Likely bugs, such as mutable defaults |
UP |
pyupgrade | Outdated syntax for your target Python version |
Two settings prevent the tools from fighting each other. First, give Ruff and Black the same line-length. Second, ignore E501 (line too long) in the linter. Black wraps what it can, and it leaves long strings and comments alone. Without the ignore, the linter fails on lines the formatter chose not to change.
If you stay on flake8, the same conflict needs this in a .flake8 file, because flake8 does not read pyproject.toml:
[flake8]
max-line-length = 88
extend-ignore = E203, W503
E203 and W503 are style rules that disagree with Black’s output for slices and line breaks around operators.
Run it automatically with pre-commit
Configuration helps only if the tools run. pre-commit runs them on changed files before each commit. Install it with python -m pip install pre-commit, then save this as .pre-commit-config.yaml.
repos:
- repo: https://github.com/charliermarsh/ruff-pre-commit
rev: v0.0.262
hooks:
- id: ruff
args: [--fix, --exit-non-zero-on-fix]
- repo: https://github.com/psf/black
rev: 23.3.0
hooks:
- id: black
pre-commit install
pre-commit run --all-files
Ruff runs before Black on purpose. A lint fix can change the code’s shape, and the formatter should have the last word. The rev values pin exact versions, so every developer and the CI server apply identical rules.
In CI, run the same tools in check mode, where they report without modifying files:
ruff check .
python -m black --check .
Both commands exit with a non-zero status when they find a problem, which fails the build. Run them before your pytest suite, because they take seconds and catch the cheapest mistakes first.
Rank lint rules by the bugs they find
Here is the insight I wish tool documentation led with. Lint rules are not equally valuable, and treating them equally is why teams come to resent linters. Sort them into three tiers.
| Tier | Rules | What a violation means | Policy |
|---|---|---|---|
| 1. Bugs | F (pyflakes), B (bugbear), E9 (syntax) |
The code is probably wrong | Always on. Never suppressed without a comment explaining why |
| 2. Maintenance | UP, I, selected simplification rules |
The code is correct but dated or untidy | On, and auto-fixed |
| 3. Taste | Naming, docstring style, complexity limits | Someone prefers it another way | Opt in one family at a time, by team agreement |
Start a legacy codebase with tier 1 only. An undefined name (F821) is a crash waiting for the right input. A naming violation is not. When the first lint run on an old project prints 4,000 style warnings, people learn to ignore the output, and the three real bugs in it go unread. A short list of true problems gets fixed.
How real systems enforce code style
- Format on save, check on commit, enforce in CI. Editors format as you type, pre-commit catches what the editor missed, and CI blocks anything that slipped past both.
- One big formatting commit, then never again. Teams adopting Black reformat the whole repository in a single commit. They list that commit in a
.git-blame-ignore-revsfile so thatgit blamestill shows the real authors. - Pinned tool versions. A new linter release can add rules and fail a build that passed yesterday. Versions are pinned and upgraded deliberately.
- Suppressions carry a code. A line-level exemption names the rule, as in
# noqa: B006. A bare# noqahides every future problem on that line. - pylint as a slower second pass. Some teams keep pylint for its deeper checks and run it in CI only, not on every commit.
In my experience moving a mid-sized service from flake8 with nine plugins plus isort to Ruff earlier this year, the pre-commit hook went from roughly twenty seconds to well under one. The migration took an afternoon, mostly spent mapping plugin codes to Ruff’s rule prefixes. The real benefit showed up later: developers stopped bypassing the hook with --no-verify, because it no longer interrupted them.
Choosing your tools: a decision framework
- Is this a new project? Use Ruff for linting and import sorting, and Black for formatting. Configure both in
pyproject.toml. - Do you have a working flake8 setup? Check whether Ruff implements the plugins you rely on. If it does, migrate. If one custom plugin is essential, keep flake8 for that plugin only.
- Do you need custom, project-specific lint rules? Ruff has no plugin system for your own rules. Use flake8 or pylint plugins for those.
- Do you want the deepest static checks? Add pylint in CI, or better, add a type checker, which finds a different and larger class of errors.
- Is the codebase old and unformatted? Adopt the formatter first in one commit, then enable tier 1 lint rules, then expand.
When NOT to use each tool
- Ruff, when you cannot tolerate churn. It is on 0.0.x releases, and rules and options still change between versions. If you need long-term stable behavior with no surprises, stay on flake8 for now, or pin Ruff tightly and upgrade on a schedule.
- Black, when you need to control the style. Black is intentionally almost unconfigurable. If a house style is a hard requirement, Black will frustrate you. Consider yapf or autopep8.
- pylint, on every commit of a large codebase. It is thorough and slow. In a pre-commit hook, it trains people to skip hooks. Run it in CI.
Common mistakes
- Running tools only by hand. Without pre-commit and CI, style drifts within weeks, and the configuration becomes decoration.
- Different line lengths per tool. The linter then rejects what the formatter produced. Every commit fails until someone aligns the numbers.
- Enabling every rule at once. Thousands of warnings teach the team to ignore all of them, including the real bugs.
- Using a bare noqa.
# noqawithout a code silences new problems added to that line later. - Mixing a reformat with a feature. A pull request that both reformats files and changes logic cannot be reviewed. Reformat in its own commit.
- Leaving tool versions unpinned. A release adds a rule overnight, and an unrelated pull request fails CI.
Key takeaways
- A linter finds problems. A formatter rewrites layout. You need both.
- Black is the standard formatter, and it has almost no options by design.
- Ruff replaces flake8, many plugins, and isort with one fast tool, and it is still on 0.0.x releases.
- Ruff does not replace Black today.
- Match line lengths, and ignore
E501in the linter when you use Black. - Run the tools from pre-commit and again in CI, with pinned versions.
- Enable bug-finding rules first, and add style rules by agreement.
FAQ
What is the difference between a linter and a formatter in Python?
A linter analyzes code and reports problems such as unused imports, undefined names, and risky patterns. A formatter rewrites whitespace, quotes, and line breaks to a consistent style without changing behavior.
Is Ruff better than flake8?
Ruff is much faster and bundles the rules of flake8, many plugins, and isort in one tool, with automatic fixes. flake8 is more mature and supports custom plugins. For new projects, Ruff is a strong default.
Does Ruff replace Black?
No. Ruff is a linter that can fix individual violations, but it does not reformat code layout. Use Black for formatting and Ruff for linting.
Should I use Black and flake8 together?
Yes, they work well together. Set flake8’s max-line-length to 88 and add extend-ignore = E203, W503 so that its style rules do not conflict with Black’s output.
Do I still need pylint if I use Ruff?
Not usually. pylint performs some deeper checks, but it is slow. Most teams get more value from Ruff plus a type checker such as mypy.
Automate the style, review the logic
Style debates end when a tool makes the decision and every commit passes through it. That frees code review for design and correctness. Pick one formatter and one linter, wire them into pre-commit and CI, and turn on the rules that find bugs first.
Rule of thumb: if a human is commenting on whitespace in a review, a tool is missing from your pipeline.
