Python GIL Explained: Threads vs Processes
Run one benchmark and see the GIL for yourself: threads speed up waiting, processes speed up computing. Learn why, and how to pick the right model.
A developer has an 8-core machine and a function that takes 10 seconds. They split the work across 8 threads and expect it to finish in a little over a second. It takes 10 seconds, sometimes 11. All 8 threads ran, the operating system scheduled them across cores, and the program still did one thing at a time.
The cause is the Python GIL, short for global interpreter lock. It is a mutex in CPython, the standard Python interpreter, that a thread must hold to execute Python bytecode. Only one thread can hold it at any moment, no matter how many cores you have.
This guide uses CPython 3.11. A thread is a line of execution inside a process that shares memory with other threads. A process is a separate running program with its own memory and its own interpreter. Run python --version before you start. Every benchmark below uses only the standard library.
My position: reach for processes before threads when the work is CPU-bound, and for threads when the work is waiting. The GIL is rarely the reason a Python program is slow, but it decides which concurrency tool can help.
See the Python GIL in one benchmark
Start with behavior you can observe. This complete script runs a CPU-bound task and an I/O-bound task in three ways. Save it as gil_demo.py.
import time
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor
WORKERS = 4
def cpu_task(n: int) -> int:
total = 0
for i in range(n):
total += i * i
return total
def io_task(seconds: float) -> float:
time.sleep(seconds) # stands in for a network or disk wait
return seconds
def timed(label: str, func) -> None:
start = time.perf_counter()
func()
print(f'{label:<16}{time.perf_counter() - start:.2f}s')
def run_sequential(task, arg) -> None:
for _ in range(WORKERS):
task(arg)
def run_pool(executor_class, task, arg) -> None:
with executor_class(max_workers=WORKERS) as pool:
list(pool.map(task, [arg] * WORKERS))
def main() -> None:
n = 5_000_000
timed('cpu sequential', lambda: run_sequential(cpu_task, n))
timed('cpu threads', lambda: run_pool(ThreadPoolExecutor, cpu_task, n))
timed('cpu processes', lambda: run_pool(ProcessPoolExecutor, cpu_task, n))
timed('io sequential', lambda: run_sequential(io_task, 1.0))
timed('io threads', lambda: run_pool(ThreadPoolExecutor, io_task, 1.0))
if __name__ == '__main__':
main()
python gil_demo.py
This is my own illustrative run on an 8-core laptop. Your numbers will differ, but the pattern will not.
cpu sequential 1.21s
cpu threads 1.24s
cpu processes 0.38s
io sequential 4.00s
io threads 1.00s
Read the two halves separately. For the CPU task, four threads were no faster than a plain loop, while four processes cut the time to about a third. For the I/O task, four threads cut four seconds to one. Same thread pool, opposite results. The rest of this article explains why.
What the GIL is and why it exists
CPython manages memory with reference counting. Every object stores a count of how many references point to it, and Python frees the object when the count reaches zero. Almost every operation changes some count.
If two threads changed the same count at once without protection, updates would be lost. Objects would be freed while still in use, or never freed. Protecting each object with its own lock would work, but it would slow down single-threaded code and risk deadlocks. Instead, CPython uses one lock around the whole interpreter.
That choice has real benefits. Single-threaded code runs fast because it takes no per-object locks. C extensions are easier to write because they can assume one thread runs at a time. The cost is the one you saw above: Python bytecode cannot execute on two cores at once inside one process.
Note the scope. The GIL belongs to CPython, not to the Python language. Other implementations make different choices. Most people run CPython, so in practice the lock applies to you.
How threads take turns
A thread does not hold the lock forever. CPython asks the running thread to release it after a time slice called the switch interval. You can read it:
python -c "import sys; print(sys.getswitchinterval())"
0.005
The default is 5 milliseconds. A waiting thread that has not received the lock within that time sets a flag. The running thread checks the flag at safe points, releases the lock, and lets another thread acquire it.
CPU-bound work, 2 threads, 1 process
Thread A [run 5ms]..........[run 5ms]..........[run 5ms]
Thread B ..........[run 5ms]..........[run 5ms]..........
only one thread holds the GIL at a time
total work per second = one core, plus switching overhead
I/O-bound work, 2 threads, 1 process
Thread A [run]--- waiting on socket (GIL released) ---[run]
Thread B [run]--- waiting on socket (GIL released) ---[run]
both waits overlap, so total time is roughly one wait
The second diagram holds the key fact. A thread releases the GIL whenever it blocks on I/O: reading a socket, waiting on a file, sleeping, or waiting for a database reply. While it waits, other threads run. That is why threads help I/O-bound programs.
C extensions can also release the lock during long computations that do not touch Python objects. NumPy does this for many array operations, as do zlib compression and hashlib on large inputs. Consequently, a threaded program that spends its time inside such calls can use several cores.
You can change the interval with sys.setswitchinterval(). A smaller value makes threads more responsive and wastes more time on switching. In practice, leave it alone. If the interval matters to you, the design probably needs processes.
Inspect a running program
To check whether the GIL limits a live process, use py-spy, a sampling profiler that attaches without code changes. Install it in a virtual environment, then point it at a process ID.
python -m venv .venv
source .venv/bin/activate # macOS and Linux
.venv\Scripts\Activate.ps1 # Windows PowerShell
python -m pip install py-spy
py-spy top --pid 12345
The header line reports how much of the time some thread held the lock:
GIL: 100.00%, Active: 100.00%, Threads: 5
A GIL figure near 100 percent, with several busy threads, means they are queueing for the lock. More threads will not help. A low figure means the threads mostly wait on I/O, and threads are doing their job. On some systems, attaching to another process needs elevated permissions.
The myth: the GIL makes threads useless
You will often read that Python threads are pointless because of the GIL. The benchmark already disproved that: threads made the I/O task four times faster. Most backend code is I/O-bound. It calls databases, HTTP services, and file systems, and it spends most of its life waiting. For that work, threads deliver nearly the full benefit.
The accurate statement is narrower: threads do not speed up CPU-bound pure-Python code. Everything else is still on the table.
The second myth: the GIL makes your code thread-safe
This one is more dangerous, because it fails silently. The GIL protects the interpreter’s internal state. It does not protect your program’s logic. A thread can lose the lock between any two steps of your code.
import threading
import time
balance = 100
def withdraw(amount: int) -> None:
global balance
if balance >= amount: # check
time.sleep(0.001) # any I/O or call here releases the GIL
balance -= amount # act
def main() -> None:
threads = [threading.Thread(target=withdraw, args=(100,)) for _ in range(2)]
for thread in threads:
thread.start()
for thread in threads:
thread.join()
print(balance)
if __name__ == '__main__':
main()
-100
Both threads passed the check before either one subtracted. The account went negative, and the GIL was held correctly the whole time. This is a race condition, and the fix is a lock around the check and the action together:
lock = threading.Lock()
def withdraw(amount: int) -> None:
global balance
with lock:
if balance >= amount:
time.sleep(0.001)
balance -= amount
With the lock, the program prints 0.
Here is an insight that older tutorials miss. The classic demonstration, two threads running counter += 1 in a loop, often shows no lost updates on recent CPython versions, because the interpreter now switches threads at fewer points. Developers run it, see the right total, and conclude that the operation is safe. It is not guaranteed. That behavior is an implementation detail that can change in any release. Never treat a passing demo as proof of thread safety. Protect every read-modify-write of shared state with a lock, or avoid shared state.
Processes: real parallelism, with costs
Each process has its own interpreter and its own GIL, so four processes can execute Python bytecode on four cores. That is why the process pool won the CPU benchmark. The concurrent.futures.ProcessPoolExecutor class gives you the same interface as the thread pool, which makes switching a one-word change.
Processes are not free, however. You pay in four ways.
| Cost | What happens | How to limit it |
|---|---|---|
| Startup | Each worker starts a new interpreter and imports your modules | Create the pool once and reuse it |
| Serialization | Arguments and results are pickled and copied between processes | Pass small inputs, such as file names or ID ranges, not large objects |
| Memory | Every worker holds its own copy of loaded data | Cap max_workers, and load data inside the worker |
| Restrictions | Functions and arguments must be picklable, so lambdas and open connections fail | Use module-level functions and plain data |
The serialization cost surprises people most. If a task takes one millisecond and its data takes two milliseconds to pickle and send, the parallel version is slower than a loop. For many small tasks, pass chunksize to pool.map() so that items travel in batches.
The benchmark sums squares with an explicit loop. A generator expression inside sum() would behave the same way under the GIL, since it is also pure Python bytecode.
One rule is mandatory. Code that starts processes must sit behind if __name__ == '__main__':. On Windows and macOS, the default start method launches a new interpreter that imports your script. Without the guard, each worker would start its own pool, and Python raises a RuntimeError about starting a process before the bootstrapping phase has finished.
Threads, processes, and asyncio compared
| Property | Threads | Processes | asyncio |
|---|---|---|---|
| Speeds up CPU-bound Python | No | Yes | No |
| Speeds up I/O-bound work | Yes | Yes, but wasteful | Yes |
| Memory per unit | Low to moderate | High | Very low |
| Shared state | Shared, so you need locks | Separate, so you pass messages | Shared, with switches only at await |
| Works with blocking libraries | Yes | Yes | No, needs async libraries |
| Practical scale | Tens to low hundreds | About one per core | Thousands of concurrent waits |
Threads and the asyncio event loop solve the same problem, waiting, in different ways. Threads let you keep ordinary blocking libraries. Asyncio scales to far more simultaneous waits but requires async code throughout.
Will the GIL go away?
The question has become active again. Earlier this month, Sam Gross published PEP 703, “Making the Global Interpreter Lock Optional in CPython”. It proposes a build option that produces an interpreter without the GIL, using techniques such as biased reference counting to keep single-threaded speed acceptable.
Treat this strictly as a proposal. It is under discussion, and no decision has been made. Open questions include the single-threaded slowdown, the work required from C extension authors, and the cost of maintaining two builds. Design your systems for the interpreter that exists today.
Meanwhile, single-threaded speed keeps improving. The faster interpreter in Python 3.11 makes each thread do more work per second, which raises the point at which you need parallelism at all.
How real systems work around the GIL
- Web servers run several worker processes. Servers such as Gunicorn and uWSGI start multiple processes, often one per core, and each handles requests independently. The GIL limits one worker, not the whole service.
- Thread pools inside each worker handle I/O. A worker uses a small thread pool to overlap database and HTTP calls for concurrent requests.
- Heavy math runs in compiled code. Numerical workloads go through NumPy and similar libraries, which do the work in C and often release the lock.
- Background jobs run in separate worker processes. Task queues such as Celery execute CPU-heavy jobs in their own processes, away from request handling.
- Process pools are created once. Long-lived services start a pool at boot and reuse it, because paying the startup cost per request would erase the gain.
In my experience profiling a thumbnail service, the team had put image resizing behind a 16-thread pool and saw no improvement over a single thread. The resize code was a pure-Python loop over pixels, so py-spy showed the GIL held nearly 100 percent of the time. Switching to a process pool helped only after we stopped passing image bytes between processes and passed file paths instead. Before that change, pickling the images took longer than resizing them.
Choosing a concurrency model: a decision framework
- Is the program actually too slow? Profile first. Concurrency adds bugs, and a better algorithm or a database index often removes the need.
- Is the bottleneck waiting or computing? If CPU usage sits near 100 percent of one core, it is computing. If the CPU is mostly idle, it is waiting.
- Waiting, with ordinary blocking libraries? Use a
ThreadPoolExecutor. Start with 10 to 30 workers and measure. - Waiting on thousands of connections at once? Use asyncio with async libraries.
- Computing in pure Python? Use a
ProcessPoolExecutorwith about one worker per core, and keep the data you pass small. - Computing on large arrays? Use NumPy or another compiled library before you add processes.
When NOT to use threads or processes
- Threads for CPU-bound pure Python. They add switching overhead and return nothing. The benchmark’s threaded CPU run was slightly slower than the sequential one.
- Processes for many tiny tasks. When each task takes microseconds, pickling and inter-process messaging cost more than the work. Batch the tasks, or stay in one process.
- Either one for a problem a library already solves. If the slow part is a nested loop over numbers, a vectorized NumPy call is usually faster than any amount of parallel Python, and much simpler.
Common mistakes
- Adding threads to CPU-bound code. The GIL serializes the work, so runtime stays flat or rises. Teams then blame the hardware.
- Assuming the GIL prevents race conditions. Check-then-act sequences on shared data still interleave. The result is corrupted state that appears only under load.
- Forgetting the main guard with processes. Without
if __name__ == '__main__':, worker startup re-runs the script on Windows and macOS and fails. - Sending large objects to workers. Pickling a big list or DataFrame for every task can take longer than the task. Parallel code ends up slower than a loop.
- Creating a pool per call. Starting processes for each request adds tens of milliseconds every time. Build the pool once.
- Using more processes than cores. Extra workers compete for the same CPUs and multiply memory use, with no throughput gain.
Key takeaways
- The GIL lets one thread at a time execute Python bytecode in a CPython process.
- Threads speed up I/O-bound work, because a waiting thread releases the lock.
- Threads do not speed up CPU-bound pure-Python work. Processes do.
- Processes cost startup time, memory, and pickling, so pass small arguments and reuse the pool.
- The GIL does not make your own code thread-safe. Guard shared state with a lock.
- Use
py-spy topto see whether threads are queueing for the lock. - PEP 703 proposes an optional no-GIL build, and it is only a proposal today.
FAQ
What is the GIL in Python?
The global interpreter lock is a mutex in CPython that allows only one thread to execute Python bytecode at a time. It protects the interpreter’s memory management, which relies on reference counting.
Does the GIL make Python single-threaded?
No. Python programs can run many threads. The GIL only prevents them from executing Python bytecode simultaneously. Threads still overlap while they wait on I/O or run C code that releases the lock.
When should I use multiprocessing instead of threading in Python?
Use multiprocessing for CPU-bound work written in pure Python, because each process has its own interpreter and lock. Use threading for I/O-bound work such as network and database calls.
Does the GIL make Python code thread-safe?
No. It protects the interpreter’s internals, not your data. Operations that read, modify, and write shared state can still interleave, so you need locks or designs that avoid shared state.
Will Python remove the GIL?
PEP 703, published in January 2023, proposes making the GIL optional through a special build of CPython. It is under discussion, and no decision has been made.
Match the tool to the bottleneck
The GIL does not make Python slow, and it does not make threads pointless. It draws one line: Python bytecode runs on one core per process. Once you know whether your program waits or computes, the choice between threads and processes follows directly.
Rule of thumb: threads for waiting, processes for computing, and a profiler before either.
