Python List Comprehension and Generators Explained
List comprehensions build collections in one readable line, and generators produce items on demand. Learn the syntax, the costs, and when to use each.
A script loads a 4 GB log file to count the lines that contain ERROR. It builds a list of every line first, and the operating system kills it for using too much memory. Changing two square brackets to parentheses makes the same script run in a few megabytes.
A Python list comprehension is an expression of the form [expression for item in iterable if condition] that creates a new list. A generator is its lazy relative: it describes the same sequence but computes each item only when something asks for it.
This guide uses Python 3.10 and only the standard library. Run python --version to check your interpreter. An iterable is any object you can loop over, such as a list, a file, or a range. If lists, sets, and dictionaries are new to you, start with Python’s built-in data structures.
My position: comprehensions and generators are tools for clarity and memory, not for speed. Choose between them by asking how many times you will walk the data and whether you need all of it at once.
Python list comprehension syntax
A comprehension replaces the common pattern of creating an empty list and appending to it in a loop. Both versions below produce the same result.
prices = [120, 45, 300, 80]
# loop version
discounted = []
for price in prices:
if price >= 100:
discounted.append(price * 0.9)
# comprehension version
discounted = [price * 0.9 for price in prices if price >= 100]
print(discounted)
[108.0, 270.0]
Read a comprehension in three parts, and it stops looking cryptic:
[ price * 0.9 for price in prices if price >= 100 ]
| | |
what to keep where items come from which items to keep (optional)
The loop variable stays inside the comprehension. After the line runs, price from the comprehension does not exist in the surrounding scope, so it cannot overwrite another variable by accident.
Filter versus transform: where the if goes
The position of if changes its meaning, and this confuses many people. An if at the end filters items out. An if ... else at the start is a conditional expression that transforms every item.
numbers = [3, -1, 4, -5]
print([n for n in numbers if n > 0]) # filter: fewer items
print([n if n > 0 else 0 for n in numbers]) # transform: same count
[3, 4]
[3, 0, 4, 0]
Dict, set, and nested forms
The same syntax builds dictionaries and sets when you use braces.
words = ['apple', 'fig', 'apple', 'kiwi']
lengths = {word: len(word) for word in words} # dict comprehension
unique_lengths = {len(word) for word in words} # set comprehension
grid = [[1, 2, 3], [4, 5, 6]]
flat = [cell for row in grid for cell in row] # nested: flatten
print(lengths)
print(unique_lengths)
print(flat)
{'apple': 5, 'fig': 3, 'kiwi': 4}
{3, 4, 5}
[1, 2, 3, 4, 5, 6]
In a nested comprehension, the for clauses run in the same order as nested loops: the outer loop comes first. However, two for clauses are the practical limit. A comprehension with three loops and two conditions saves lines and costs every later reader several minutes.
Generator expressions: the lazy version
Swap the square brackets for parentheses and you get a generator expression. It does not build a list. Instead, it returns a generator object that computes one item each time you ask.
squares_list = [n * n for n in range(5)]
squares_gen = (n * n for n in range(5))
print(squares_list)
print(squares_gen)
print(next(squares_gen))
print(next(squares_gen))
print(list(squares_gen))
[0, 1, 4, 9, 16]
<generator object <genexpr> at 0x7f3c2a1b4040>
0
1
[4, 9, 16]
The built-in next() pulls a single item. A for loop, sum(), or list() pulls items until the generator runs out. When a generator expression is the only argument to a function, you can drop the extra parentheses: sum(n * n for n in range(5)).
Measure the memory difference
The reason to care is memory. This is my own illustrative test on Python 3.10, and you can reproduce it by saving the file and running it.
import sys
def main() -> None:
squares_list = [n * n for n in range(1_000_000)]
squares_gen = (n * n for n in range(1_000_000))
print(f'list: {sys.getsizeof(squares_list):,} bytes')
print(f'generator: {sys.getsizeof(squares_gen):,} bytes')
if __name__ == '__main__':
main()
list: 8,448,728 bytes
generator: 104 bytes
The list holds a million references, about 8.4 MB, and that figure excludes the integer objects themselves. The generator holds only its current position, so its size stays the same for a billion items. The trade-off is that a generator cannot tell you its length, cannot be indexed, and cannot be rewound.
A generator runs once
After a generator produces its last item, it is exhausted. A second pass yields nothing, and Python raises no error.
totals = (amount for amount in [10, 20, 30])
print(sum(totals)) # 60
print(sum(totals)) # 0, silently
This silent zero is the most common generator bug. If you need the data twice, build a list.
Generator functions with yield
A generator expression fits one line. For anything longer, write a generator function: a normal function that uses yield instead of return. Calling it returns a generator and runs no code yet. Each next() runs the body until the next yield, then pauses with its local variables intact.
def countdown(start: int):
print('starting')
while start > 0:
yield start
start -= 1
print('done')
timer = countdown(2)
print('created')
for value in timer:
print(value)
created
starting
2
1
done
Notice that created prints before starting. The function body does not begin until the loop asks for the first value. If you want a refresher on ordinary calls first, read about how functions take arguments and return values.
Build a lazy pipeline
Generators compose. Each stage pulls one item from the stage before it, so the whole chain holds one item in memory at a time. Save this complete example as pipeline.py.
import sys
from itertools import islice
from pathlib import Path
from typing import Iterable, Iterator
def read_lines(path: Path) -> Iterator[str]:
with path.open(encoding='utf-8') as handle:
for line in handle:
yield line.rstrip('\n')
def errors_only(lines: Iterable[str]) -> Iterator[str]:
return (line for line in lines if ' ERROR ' in line)
def messages(lines: Iterable[str]) -> Iterator[str]:
for line in lines:
yield line.split(' ERROR ', 1)[1]
def main() -> int:
if len(sys.argv) != 2:
print('usage: python pipeline.py LOG_FILE')
return 2
pipeline = messages(errors_only(read_lines(Path(sys.argv[1]))))
for message in islice(pipeline, 3):
print(message)
return 0
if __name__ == '__main__':
raise SystemExit(main())
Create a file named app.log with these lines:
10:00:01 INFO started
10:00:05 ERROR database timeout
10:00:09 INFO retrying
10:00:12 ERROR database timeout
10:00:20 ERROR disk full
10:00:21 ERROR disk full
python pipeline.py app.log
database timeout
database timeout
disk full
itertools.islice takes the first three results and stops. Consequently, the script never reads the sixth line. On a multi-gigabyte file, it would stop after the third match and use a few kilobytes. The cost is debuggability: you cannot print a generator to see its contents, and an error surfaces at the point of consumption, not where you built the stage.
Delegate with yield from
yield from passes through every item of another iterable. It keeps recursive generators short.
def flatten(items):
for item in items:
if isinstance(item, list):
yield from flatten(item)
else:
yield item
print(list(flatten([1, [2, [3, 4]], 5])))
[1, 2, 3, 4, 5]
The myth: comprehensions and generators are always faster
You will often read that a comprehension is “the fast way” and that a generator is faster still. The first claim is weaker than it sounds, and the second is usually wrong.
A list comprehension is modestly faster than an equivalent append loop, because it skips the repeated method lookup. You can check with my illustrative test:
python -m timeit "result = []" "for n in range(100_000): result.append(n * 2)"
python -m timeit "result = [n * 2 for n in range(100_000)]"
The gap is real but small, and it disappears next to any I/O. A generator, by contrast, is often slightly slower per item than a list, because Python resumes a paused frame for each value. A generator wins on memory, and on time only when you stop early.
That last case is the insight worth remembering. Functions such as any() and all() stop at the first decisive item, but only if you let them be lazy:
python -m timeit "any([n > 10 for n in range(1_000_000)])"
python -m timeit "any(n > 10 for n in range(1_000_000))"
The first command builds a million-item list before any() looks at it. The second stops after twelve items. On my machine the difference is several orders of magnitude, and you can reproduce it with those two commands.
There is also a counterexample that surprises people. str.join() needs to walk its input twice, so CPython converts a generator argument into a list internally. As a result, ', '.join([str(n) for n in items]) is slightly faster than the generator form. In short, laziness helps the consumers that can stop early or stream, and does nothing for the ones that need everything.
List comprehension versus generator: a comparison
| Property | List comprehension | Generator expression or function |
|---|---|---|
| Syntax | [x for x in items] |
(x for x in items) or yield |
| When work happens | Immediately, all at once | On demand, one item at a time |
| Memory | Grows with the number of items | Constant |
| Iterate more than once | Yes | No, exhausted after one pass |
len(), indexing, slicing |
Yes | No (use itertools.islice) |
| Infinite sequences | Impossible | Fine |
| Debugging | Print it and inspect | Must consume it to see values |
How real systems use comprehensions and generators
Production code uses each form in recognizable places.
- Comprehensions for reshaping small data. Code turns database rows into response dictionaries, or builds an index such as
{user.id: user for user in users}. The data fits in memory and gets used more than once. - Generators for file and network streams. Log processors, CSV importers, and export jobs read one record at a time. Memory stays flat whether the input has a thousand rows or a billion.
- Generators for pagination. A function loops over API pages and yields each record. The caller writes one
forloop and never sees the page boundaries. - Batching with islice. Bulk database inserts pull fixed-size chunks from a generator with
itertools.islice, so each insert statement stays a manageable size. - Generator arguments to reducers. Calls such as
sum(order.total for order in orders)andmax(len(name) for name in names)avoid a temporary list.
We once hit a bug when a report function received a generator of invoices, logged len(list(invoices)) for debugging, and then looped over invoices to compute totals. The debug line consumed the generator, so every report showed zero revenue. Nothing raised an error. Since then, I convert to a list at the boundary whenever a function reads its input more than once.
Choosing between a loop, a comprehension, and a generator: a decision framework
Ask these questions in order and stop at the first clear answer.
- Are you running code for its side effects? If the body prints, writes, or sends something, use a plain
forloop. A comprehension exists to build a value. - Does the logic need more than two clauses, or a try block? Use a loop, or move the logic into a generator function. Readability beats brevity.
- Could the data be large, unbounded, or streamed? Use a generator, so memory does not grow with the input.
- Will you pass the result straight into one consumer? For
sum,any,all,max, or a singleforloop, use a generator expression. - Do you need the length, an index, or a second pass? Use a list comprehension.
When NOT to use a comprehension or a generator
- When you only want the side effect.
[print(item) for item in items]builds and throws away a list ofNonevalues. It also tells readers that a result matters when it does not. - When the expression needs error handling. You cannot put
tryandexceptinside a comprehension. If one bad record should be skipped and logged, write a loop. - When a generator would outlive its resource. A generator that reads from an open file or database cursor keeps that resource open until it finishes. If you return it from a function and the caller never consumes it, the handle stays open. In that case, return a list.
Common mistakes
- Consuming a generator twice. The second pass produces nothing and raises no error. Totals come out as zero and nobody notices until a report looks wrong.
- Wrapping a generator in a list for any() or all(). The brackets force the whole sequence to be built first. You lose the early exit and pay the full memory cost.
- Confusing the filter if with the conditional expression. Putting
elseafter a trailingifis aSyntaxError. The conditional form goes beforefor. - Nesting three or more levels. The line passes review because it is short, then nobody can modify it safely. A later bug fix turns into a rewrite.
- Calling len() on a generator. Python raises
TypeError: object of type 'generator' has no len(). Counting withsum(1 for _ in gen)works but exhausts it. - Building a huge list to iterate once.
for line in [l.strip() for l in handle]loads the entire file. Memory use scales with file size for no benefit.
Key takeaways
- Read a comprehension as three parts: what to keep, where it comes from, and which items pass.
- A trailing
iffilters, and a leadingif ... elsetransforms. - Square brackets build everything now. Parentheses compute one item at a time.
- A generator uses constant memory but supports only one pass, with no length and no indexing.
- Pass generator expressions directly to
any,all,sum, andmaxto avoid temporary lists. - Chain generator functions to process files of any size with flat memory.
- Switch to a plain loop when you need side effects, error handling, or more than two clauses.
FAQ
What is a list comprehension in Python?
It is an expression that builds a new list from an iterable in one line, in the form [expression for item in iterable if condition]. It replaces an empty list plus an append loop.
What is the difference between a list comprehension and a generator expression?
A list comprehension uses square brackets and builds the entire list in memory at once. A generator expression uses parentheses and produces items one at a time, so it uses constant memory but can be iterated only once.
Are list comprehensions faster than for loops?
Slightly, when the loop only appends to a list. The difference is small and rarely matters next to I/O. Choose a comprehension for clarity, not speed.
What does yield do in Python?
yield turns a function into a generator. Each time the caller asks for a value, the function runs to the next yield, hands back that value, and pauses with its local state preserved.
Can I reuse a generator in Python?
No. A generator is exhausted after one pass. Create a new one by calling the generator function again, or store the results in a list if you need several passes.
Brackets when you keep it, parentheses when you pass it on
Comprehensions and generators express the same idea with different costs. One pays in memory for the freedom to reuse the result. The other pays in flexibility for flat memory and early exits. Decide by how the data gets consumed, not by which looks more clever.
Rule of thumb: if you will touch the data once, generate it; if you will touch it twice, list it; if it needs a comment to explain, loop it.
