How to Profile and Speed Up Python Code: cProfile, py-spy, lru_cache
Find the slow line before you change anything. Use cProfile, line_profiler, and py-spy to locate the hotspot, then fix it with the right technique.
cProfile counts every function call, line_profiler times individual lines, and py-spy samples a running process with almost no overhead. Profiling costs a few minutes, and it prevents days spent optimizing code that was never slow. Find the single largest hotspot, fix it with a better data structure or a cache, and measure again.A team spent a week rewriting a JSON serializer because “serialization is always the bottleneck”. The endpoint got 3 percent faster. A ten-minute profile afterwards showed that 80 percent of each request went to a list membership test inside a loop. Changing one list to a set took a single line and cut the response time by more than half.
To profile Python code means to run it under a tool that records how long each function or line takes and how often it runs. The result is a ranked list of where time goes. The parts at the top are called hotspots.
This guide uses Python 3.11. cProfile, pstats, timeit, and tracemalloc ship with Python. The third-party tools, py-spy, line_profiler, and snakeviz, install with pip. Run python --version first.
My position: measure before you optimize, because your guess is usually wrong. I have profiled many slow services, and the hotspot was where the team expected perhaps one time in five.
A slow script to work on
You need something to measure. This script builds 20,000 order IDs, cleans them, and reports the duplicates. Save it as report.py.
import re
import time
def make_orders(count: int) -> list[str]:
return [f'ord-{i % (count // 2):06d}' for i in range(count)]
def normalize(order: str) -> str:
return re.sub(r'[^A-Z0-9]', '', order.upper())
def find_duplicates(orders: list[str]) -> list[str]:
seen: list[str] = []
duplicates: list[str] = []
for order in orders:
if order in seen:
duplicates.append(order)
else:
seen.append(order)
return duplicates
def main() -> None:
start = time.perf_counter()
orders = [normalize(order) for order in make_orders(20_000)]
duplicates = find_duplicates(orders)
elapsed = time.perf_counter() - start
print(f'{len(duplicates)} duplicates in {elapsed:.2f}s')
if __name__ == '__main__':
main()
python report.py
10000 duplicates in 2.31s
The timing is from my laptop, and yours will differ. Before you read on, guess which function is slow. Many readers pick normalize, because it runs a regular expression 20,000 times.
Profile Python code with cProfile
cProfile is the standard library’s deterministic profiler. It records every function call and return. Run it from the command line and sort by cumulative time:
python -m cProfile -s cumulative report.py
Trimmed output from my machine:
ncalls tottime percall cumtime percall filename:lineno(function)
1 0.000 0.000 2.412 2.412 report.py:1(<module>)
1 0.001 0.001 2.411 2.411 report.py:24(main)
1 2.335 2.335 2.337 2.337 report.py:13(find_duplicates)
1 0.006 0.006 0.069 0.069 report.py:26(<listcomp>)
20000 0.011 0.000 0.063 0.000 report.py:9(normalize)
20000 0.012 0.000 0.046 0.000 re.py:197(sub)
Three columns tell the story.
| Column | Meaning | Use it to find |
|---|---|---|
ncalls |
How many times the function was called | Functions called far more often than you expected |
tottime |
Time spent inside the function itself, excluding the functions it calls | The function whose own code is slow |
cumtime |
Time in the function plus everything it calls | The slow branch of the call tree |
Read cumtime from the top to follow the slow path downward. Then look for the row where tottime is high: that is where the work actually happens. Here, find_duplicates owns 2.3 of 2.4 seconds in its own body. normalize, the popular guess, accounts for under 3 percent.
For larger programs, save the data and explore it:
python -m cProfile -o report.prof report.py
python -m pip install snakeviz
snakeviz report.prof
snakeviz opens an interactive chart in your browser, where the widest block is the biggest cost. To profile only part of a program, use the profiler as a context manager:
import cProfile
import pstats
with cProfile.Profile() as profiler:
find_duplicates(orders)
pstats.Stats(profiler).sort_stats('tottime').print_stats(5)
Find the slow line with line_profiler
cProfile stops at the function level. It says find_duplicates is slow, but not which line. line_profiler answers that. Install it, then mark the function with @profile. The decorator needs no import, because the kernprof runner injects it.
python -m venv .venv
source .venv/bin/activate # macOS and Linux
.venv\Scripts\Activate.ps1 # Windows PowerShell
python -m pip install line_profiler py-spy
kernprof -l -v report.py
Trimmed output:
Line # Hits % Time Line Contents
==========================================
13 @profile
14 def find_duplicates(orders):
17 20000 0.1 for order in orders:
18 20000 99.6 if order in seen:
19 10000 0.1 duplicates.append(order)
21 10000 0.2 seen.append(order)
One line takes 99.6 percent of the function. order in seen scans a list, and the list grows to 10,000 items. The loop therefore performs about 100 million comparisons. Remove the @profile line when you are done, since plain Python does not define that name.
Fix the hotspot, then measure again
A set tests membership in constant time on average, because it uses a hash table instead of a scan.
# WRONG: list membership inside a loop is a hidden nested loop
seen: list[str] = []
if order in seen: ...
# RIGHT: set membership is constant time on average
def find_duplicates(orders: list[str]) -> list[str]:
seen: set[str] = set()
duplicates: list[str] = []
for order in orders:
if order in seen:
duplicates.append(order)
else:
seen.add(order)
return duplicates
python report.py
10000 duplicates in 0.07s
On my machine, the run drops from 2.31 seconds to 0.07 seconds, with a three-line change. Reproduce it by editing the function and running the script again.
Now profile again. This step matters, because the ranking has changed. normalize and re.sub are now the top entries, simply because the giant is gone. You might precompile the pattern with re.compile(). Measure it, and you will find a small gain, since the re module already caches compiled patterns. At 0.07 seconds, the honest decision is to stop.
The original insight: check the ceiling before you start
Before you optimize a function, compute the best possible outcome. If a function takes a fraction p of total time, making it infinitely fast gives a maximum speedup of 1 / (1 - p). This is Amdahl’s law, and it takes ten seconds to apply to a profile.
| Share of total time | Best possible speedup | Worth the effort? |
|---|---|---|
| 5 percent | 1.05x | No |
| 20 percent | 1.25x | Rarely |
| 50 percent | 2x | Often |
| 90 percent | 10x | Yes |
| 97 percent | 33x | Yes, do this first |
In the example, find_duplicates held about 97 percent, so a large win was available, and we got it. normalize held 3 percent, so even a perfect rewrite could not have reached 1.05x. My rule: if the top function holds less than a quarter of the time, no single fix will transform the program. You are then looking at a design change, not a tuning job.
Profile a running process with py-spy
cProfile requires you to start the program under the profiler. Production processes are already running, and you cannot restart them casually. py-spy is a sampling profiler. It reads the process’s call stack from outside, many times per second, without modifying your code.
py-spy top --pid 12345 # live view, like top
py-spy record -o profile.svg --pid 12345 # flame graph of a running process
py-spy record -o profile.svg -- python report.py # launch and record
py-spy dump --pid 12345 # print every thread's stack once
py-spy record writes a flame graph. Each bar is a function, and its width is the share of samples in which that function was on the stack. Look for wide bars near the top of the stacks. py-spy dump is the fastest way to learn what a hung process is doing: it prints the current stack of every thread and exits. On some systems, attaching to another process needs elevated permissions.
Deterministic versus sampling profilers
| Property | cProfile (deterministic) | py-spy (sampling) |
|---|---|---|
| How it works | Hooks every function call and return | Reads the stack about 100 times per second |
| Overhead | High for code with many small calls | Very low |
| Exact call counts | Yes | No |
| Attach to a running process | No | Yes |
| Safe in production | Rarely | Generally yes |
| Sees time in C extensions | As one opaque call | Yes, with the --native option |
One caveat follows from the first row. cProfile adds a fixed cost to every call, so it exaggerates the weight of code made of many tiny functions. If a profile blames a trivial helper called a million times, confirm with wall-clock timing or a sampling profiler before you inline anything.
Two other tools are worth knowing. scalene profiles CPU and memory together and separates Python time from native time. timeit, from the standard library, compares two small snippets fairly:
python -m timeit -s "data = list(range(10_000))" "9_999 in data"
python -m timeit -s "data = set(range(10_000))" "9_999 in data"
Speed up repeated work with lru_cache
After data structures, the next most common fix is to stop repeating work. functools.lru_cache remembers a function’s results by its arguments. Its simpler sibling, functools.cache, does the same with no size limit.
import functools
import time
@functools.lru_cache(maxsize=256)
def tax_rate(region: str) -> float:
time.sleep(0.2) # stands in for a slow lookup
return {'north': 0.2, 'south': 0.07}.get(region, 0.0)
def main() -> None:
start = time.perf_counter()
regions = ['north', 'south'] * 50
total = sum(100 * tax_rate(region) for region in regions)
print(f'total {total:.0f} in {time.perf_counter() - start:.2f}s')
print(tax_rate.cache_info())
if __name__ == '__main__':
main()
total 1350 in 0.40s
CacheInfo(hits=98, misses=2, maxsize=256, currsize=2)
Without the cache, 100 lookups at 0.2 seconds each would take 20 seconds. With it, only the two distinct regions pay the cost. cache_info() proves the cache is working: 98 hits and 2 misses.
A cache trades memory and freshness for speed, so apply it with care:
- Only cache pure functions. The result must depend on the arguments alone. A cached function that reads the clock or a database returns stale answers forever.
- Arguments must be hashable. Passing a list raises
TypeError: unhashable type: 'list'. - Set a maxsize. An unbounded cache keyed by user input is a memory leak.
- Avoid it on instance methods. The cache holds a reference to
self, so the objects are never freed.
Profile memory with tracemalloc
Sometimes the problem is memory, not time. tracemalloc records where allocations happen.
import tracemalloc
tracemalloc.start()
orders = [f'ord-{i:06d}' for i in range(200_000)]
snapshot = tracemalloc.take_snapshot()
for stat in snapshot.statistics('lineno')[:3]:
print(stat)
Each line of output names a file, a line number, the total size, and the number of allocations. When a list of millions of items dominates, consider whether you need them all in memory at once. Streaming the data with generators often removes the peak entirely.
The myth: micro-optimizations make Python fast
Advice lists are full of tricks: bind a method to a local variable, replace + with join, or avoid dots in loops. These tips are popular because they are easy to apply without measuring. Each can shave a few percent from a hot loop. None of them fixes a program that does the wrong amount of work.
The example made this concrete. No micro-optimization of the list version comes anywhere near the set version, because the difference is 100 million comparisons against 20,000 hash lookups. Fixes come in a rough order of payoff:
- Do less work. Choose a better algorithm or data structure.
- Do it once. Cache results, and move repeated work out of loops.
- Do it in batches. Replace many small queries or requests with one large one.
- Do it in compiled code. Use built-ins and libraries such as NumPy for heavy loops.
- Do it in parallel. Use more cores, which for CPU-bound Python means processes. See how the GIL and process pools interact.
- Tune the remaining hot loop. Micro-optimizations belong here, last.
Upgrading the interpreter sits outside that list, because it needs no code change. Python 3.11’s faster interpreter speeds up pure-Python code as it stands.
How real systems find and fix slow code
- Metrics first, profiler second. Request timing and tracing show which endpoint or job is slow. A profiler then explains why. Profiling a whole service without a target wastes time.
- Sampling in production. Teams attach
py-spyto a live worker for 30 seconds during a slow period, and read the flame graph. - The database is the usual suspect. In web services, a profile most often points at many small queries, not at Python code. The fix is fewer, larger queries.
- Benchmarks guard the fix. After a hotspot is fixed, a small timing test or a query-count assertion keeps it from returning.
- Caches get limits and expiry. In-process caches have a maximum size, and anything that can change has a time limit or an explicit invalidation path.
In my experience profiling a nightly export that had grown from four minutes to over an hour, everyone suspected the CSV writing. A py-spy flame graph showed almost all samples under one line that checked each row’s ID against a list of already exported IDs. The list had grown with the customer base, so the job’s cost grew with the square of the data. Replacing the list with a set brought the run back to about three minutes.
Choosing a profiling tool: a decision framework
- Can you reproduce the slowness in a script or test? Start with
python -m cProfile -s cumulative. - Is the process already running, or in production? Use
py-spy toporpy-spy record. - Do you know the slow function but not the line? Use
line_profileron that function. - Is the process stuck? Use
py-spy dumpto see every thread’s stack. - Is memory the problem? Use
tracemallocorscalene. - Are you comparing two ways to write one expression? Use
timeit.
When NOT to optimize
- When the code is fast enough. A nightly job that takes two minutes and has an hour to run needs no work. Optimized code is usually harder to read.
- When the time is spent waiting. If a profile shows the process idle in network or database calls, faster Python changes nothing. Fix the query, add an index, or overlap the waits.
- When the hotspot’s share is small. A function that holds 5 percent of the time cannot give you more than a 1.05x speedup. Leave it alone.
Common mistakes
- Optimizing without a profile. Effort goes to code that looks slow instead of code that is slow. The program barely changes.
- Profiling with toy data. A quadratic loop is invisible at 100 rows and catastrophic at 100,000. Use production-sized inputs.
- Reading only cumtime. The top rows are always
mainand its callers. Withouttottime, you never reach the function doing the work. - Trusting cProfile timings as absolute. Profiler overhead inflates call-heavy code. Confirm the improvement with wall-clock time.
- Caching an impure function. A cached lookup keeps returning old data after the source changes. Users see stale prices or permissions.
- Not measuring after the fix. Some “optimizations” make code slower. Without a second measurement, they ship anyway.
Key takeaways
- Profile first. The slow part is rarely where you expect.
- Run
python -m cProfile -s cumulative script.py, and look for hightottime. - Use
line_profilerto find the exact line inside a slow function. - Use
py-spyfor running and production processes. - Compute the ceiling: a function with share
pcan give at most1 / (1 - p). - Fix algorithms and data structures first, then cache, then batch.
- Measure again after every change, and stop when the code is fast enough.
FAQ
How do I profile Python code?
Run your script with python -m cProfile -s cumulative script.py. The output lists each function with its call count and time. Look for functions with a high tottime value.
What is the difference between tottime and cumtime in cProfile?
tottime is the time spent in a function’s own code, excluding the functions it calls. cumtime includes the time of everything it calls.
What is py-spy used for?
py-spy is a sampling profiler that attaches to a running Python process without code changes or a restart. It shows a live view of hot functions, records flame graphs, and dumps thread stacks.
When should I use lru_cache in Python?
Use it on pure functions that are expensive and called repeatedly with the same hashable arguments. Set a maxsize so that the cache cannot grow without limit.
Does cProfile slow down my program?
Yes. cProfile adds overhead to every function call, so programs with many small calls can run noticeably slower under it. Use a sampling profiler such as py-spy when overhead matters.
Measure first, fix the biggest thing, measure again
Performance work is a loop with three steps, and skipping the first one is the common failure. A profile turns an argument about what might be slow into a ranked list. Fix the top entry with the cheapest technique that works, confirm the gain, and stop when you have enough.
Rule of thumb: never optimize a line you have not seen at the top of a profile.
