Python Data Structures: Lists, Tuples, Sets, and Dictionaries Explained

Lists, tuples, sets, and dictionaries each answer a different question about your data. Learn what each costs, when to pick it, and the bugs to avoid.

Executive Summary: The four built-in Python data structures are the list, the tuple, the set, and the dictionary. Each one is fast at a different job: lists keep order, tuples hold fixed records, sets test membership, and dictionaries look up values by key. Choosing the wrong one rarely breaks your code, but it often makes it slow or fragile. Therefore, pick the structure by the question you ask of your data most often.

A script checks 100,000 incoming IDs against 100,000 known IDs. With the known IDs in a list, it runs for minutes. With the same IDs in a set, it finishes in a fraction of a second. The code differs by one pair of brackets.

Python data structures are the built-in container types that hold groups of values: list, tuple, set, and dict. They differ in three ways: whether they keep order, whether you can change them, and how fast they find an item.

This guide uses Python 3.10 and only the standard library, so you can paste every example into a file and run it. Check your interpreter with python --version first. If you later install third-party packages, create a virtual environment with venv so that each project keeps its own packages.

My position: most slow or buggy beginner code uses a list where a set or a dictionary belongs. Learn the four questions below and you will choose correctly almost every time.

The four Python data structures at a glance

Start with a side-by-side view. The table shows the properties that drive every later decision.

Type Literal Ordered Mutable Duplicates Fast at
list [1, 2, 3] Yes Yes Allowed Access by position, append at the end
tuple (1, 2, 3) Yes No Allowed Fixed records, dictionary keys
set {1, 2, 3} No Yes Removed Membership tests, removing duplicates
dict {'a': 1} Insertion order Yes Keys unique Lookup by key

Here is the heuristic I teach: choose by the question you ask most often.

  • “What is at position i?” Use a list.
  • “Which fields make up this one record?” Use a tuple.
  • “Have I seen this value before?” Use a set.
  • “What belongs to this key?” Use a dictionary.

Lists: ordered and changeable

A list is an ordered sequence that you can grow, shrink, and reorder. It is the default container in Python, and it suits any collection where position matters.

tasks = ['write tests', 'fix bug', 'deploy']

tasks.append('update docs')      # add at the end
first = tasks[0]                 # read by position
last = tasks[-1]                 # negative index counts from the end
urgent = tasks[:2]               # slice: a new list with the first two items
tasks.sort()                     # sort in place

print(first, last)
print(urgent)
print(tasks)
write tests update docs
['write tests', 'fix bug']
['deploy', 'fix bug', 'update docs', 'write tests']

Note the difference between tasks.sort() and sorted(tasks). The method changes the list and returns None. The function leaves the list alone and returns a new one. Writing tasks = tasks.sort() therefore replaces your list with None.

Lists are fast at the end and slow at the front. append and pop() take constant time. However, insert(0, x) and pop(0) shift every remaining item, so their cost grows with the length of the list.

Tuples: fixed records

A tuple is an ordered sequence that you cannot change after you create it. Use it for a small group of related values where each position has a meaning, such as a coordinate or a database row.

point = (12.5, 48.1)
latitude, longitude = point      # unpacking

def min_max(values):
    return min(values), max(values)   # returns a tuple

low, high = min_max([4, 9, 1, 7])
print(latitude, longitude, low, high)
12.5 48.1 1 9

Two syntax details trip people up. First, the comma makes a tuple, not the parentheses, so a one-item tuple needs a trailing comma: (5,). Second, trying to change a tuple raises TypeError: 'tuple' object does not support item assignment.

Immutability has a practical payoff. Because a tuple of immutable values cannot change, Python can hash it, so you can use it as a dictionary key or put it in a set. A list cannot do either job.

List versus tuple

People often describe a tuple as a list you cannot change. That description misses the point. A list holds many items of the same kind, and its length varies. A tuple holds one record with a fixed shape. If you would give each position a name, you want a tuple.

Sets: unique values and fast membership

A set is an unordered collection of unique values. It answers one question very quickly: is this value present? It also supports the set operations you know from mathematics.

backend = {'ana', 'ben', 'chen'}
on_call = {'ben', 'dara'}

print('ben' in backend)          # membership
print(backend & on_call)         # intersection: in both
print(backend | on_call)         # union: in either
print(backend - on_call)         # difference: in backend only

emails = ['a@example.com', 'b@example.com', 'a@example.com']
print(len(set(emails)))          # remove duplicates
True
{'ben'}
{'ana', 'ben', 'chen', 'dara'}
{'ana', 'chen'}
2

The order of items in your printed sets may differ, because a set has no order. Do not write code that depends on it. Also remember that {} creates an empty dictionary. An empty set is set().

The trade-off is memory and restrictions. A set uses more memory than a list of the same items, and it accepts only hashable values. Adding a list to a set raises TypeError: unhashable type: 'list'.

Dictionaries: lookup by key

A dictionary maps keys to values. It is the structure behind JSON objects, configuration, counters, caches, and most real Python programs.

stock = {'keyboard': 14, 'mouse': 30}

stock['monitor'] = 5                     # add or replace
print(stock['mouse'])                    # read: raises KeyError if missing
print(stock.get('webcam', 0))            # read with a default

for product, quantity in stock.items():  # iterate over pairs
    print(product, quantity)

defaults = {'theme': 'light', 'page_size': 20}
overrides = {'page_size': 50}
settings = defaults | overrides          # merge: right side wins
print(settings)
30
0
keyboard 14
mouse 30
monitor 5
{'theme': 'light', 'page_size': 50}

Use square brackets when a missing key is a bug, because the KeyError tells you early. Use get when a missing key is normal. The | merge operator arrived in Python 3.9, so it fails with a TypeError on older interpreters.

The myth: dictionaries are unordered

Many tutorials and interview answers still say that a dictionary has no order. That was true years ago, and it is now wrong. Since Python 3.7, the language guarantees that a dictionary keeps insertion order. Iteration returns keys in the order you added them.

However, do not confuse insertion order with sorted order. A dictionary does not sort its keys, and you cannot ask for the item at position 3. If you need sorted output, call sorted(stock). Sets, by contrast, remain unordered.

What each operation costs

Speed differences between these types are large, and they come from how each type stores data. A list is a row of slots that Python scans one by one. A set or dictionary uses a hash table: Python computes a number from the value and jumps straight to the right slot.

Is 'dara' in the collection?

list:  ['ana', 'ben', 'chen', 'dara']
         no     no     no      yes        4 comparisons, grows with length

set:   hash('dara') -> slot 2 -> found     1 jump, same for any size
Operation list tuple set dict
Read by position Constant Constant Not supported Not supported
Membership test (in) Grows with size Grows with size Constant on average Constant on average (keys)
Add at the end Constant Not supported Constant on average Constant on average
Insert or remove at the front Grows with size Not supported Not applicable Not applicable
Lookup by key Not supported Not supported Not applicable Constant on average

You can measure the gap yourself. This is my own illustrative test, and your numbers will vary by machine. Each command looks for the last item among 100,000 integers.

python -m timeit -s "data = list(range(100_000))" "99_999 in data"
python -m timeit -s "data = set(range(100_000))" "99_999 in data"

On my laptop, the list lookup takes around a millisecond and the set lookup takes tens of nanoseconds. That is a difference of several orders of magnitude for one check. Put the check inside a loop of 100,000 items, and the list version does billions of comparisons.

The original insight here is a code smell you can search for. An in test against a list, written inside a loop, is a hidden nested loop. Whenever you see that pattern, convert the list to a set before the loop starts.

Copying and aliasing: the bug that surprises everyone

Assignment never copies a container in Python. It gives the same object a second name. Changes made through one name appear through the other.

# WRONG: both names point at one list
original = [1, 2, 3]
backup = original
backup.append(4)
print(original)        # [1, 2, 3, 4]
# RIGHT: make a real copy
original = [1, 2, 3]
backup = original.copy()
backup.append(4)
print(original)        # [1, 2, 3]

The copy() method makes a shallow copy: a new outer container that still shares its inner objects. For nested data, you need a deep copy, which duplicates every level.

import copy

grid = [[0, 0], [0, 0]]

shallow = grid.copy()
shallow[0][0] = 9
print(grid)            # [[9, 0], [0, 0]]  inner lists are shared

grid = [[0, 0], [0, 0]]
deep = copy.deepcopy(grid)
deep[0][0] = 9
print(grid)            # [[0, 0], [0, 0]]  fully independent

The same sharing explains a classic trap when you build a grid:

# WRONG: three references to the same row
board = [[0] * 3] * 3
board[0][0] = 1
print(board)           # [[1, 0, 0], [1, 0, 0], [1, 0, 0]]

# RIGHT: a new row on each pass
board = [[0] * 3 for _ in range(3)]
board[0][0] = 1
print(board)           # [[1, 0, 0], [0, 0, 0], [0, 0, 0]]

To check whether two names share one object, use is. The == operator compares contents, while is compares identity. Two separate lists with equal items are == but not is.

Deep copies are safe but slow on large structures. In practice, copy only at the boundary where you hand data to code you do not control.

The collections module: specialized containers

The standard library’s collections module adds containers for jobs the four basic types handle awkwardly. Four of them cover most needs.

Container Replaces Use it for
Counter A dict of counts built by hand Counting items and finding the most common
defaultdict “If key missing, create it” checks Grouping values under keys
deque A list used as a queue Fast appends and pops at both ends
namedtuple A plain tuple with magic positions Small records with named fields

Here is a complete program that uses a tuple, a set, a Counter, and a defaultdict together. Save it as orders.py.

from collections import Counter, defaultdict

ORDERS = [
    ('alice', 'keyboard', 2),
    ('bob', 'mouse', 1),
    ('alice', 'mouse', 1),
    ('carol', 'monitor', 1),
    ('bob', 'keyboard', 1),
]


def summarize(orders: list[tuple[str, str, int]]):
    customers = set()
    units_by_product = Counter()
    products_by_customer = defaultdict(list)
    for customer, product, quantity in orders:
        customers.add(customer)
        units_by_product[product] += quantity
        products_by_customer[customer].append(product)
    return customers, units_by_product, products_by_customer


def main() -> None:
    customers, units, by_customer = summarize(ORDERS)
    print(f'customers: {sorted(customers)}')
    for product, total in units.most_common():
        print(f'{product}: {total}')
    for customer, products in by_customer.items():
        print(f'{customer} bought {products}')


if __name__ == '__main__':
    main()
python orders.py
customers: ['alice', 'bob', 'carol']
keyboard: 3
mouse: 2
monitor: 1
alice bought ['keyboard', 'mouse']
bob bought ['mouse', 'keyboard']
carol bought ['monitor']

Each order is a tuple because it is one fixed record. The customers go in a set because only uniqueness matters. The counts and groups use dictionary subclasses that create missing keys for you. For a queue, use deque: its popleft() takes constant time, while list.pop(0) slows down as the list grows.

How real systems use these structures

Production Python code leans on a few recurring patterns. Recognizing them helps you read other people’s code quickly.

  • Dictionaries as indexes. Code loads rows from a database or file once, then builds a dictionary keyed by ID. Every later lookup takes constant time instead of scanning the rows again.
  • Sets for “already seen” tracking. Crawlers, deduplication jobs, and idempotency checks keep a set of processed IDs and skip repeats.
  • Tuples as composite keys. A cache keyed by (user_id, date) uses a tuple, because the pair identifies one entry and tuples are hashable.
  • Lists for ordered results. API responses, query results, and batches stay in lists, since their order matters and code walks them from start to end.
  • Deques for buffers. A deque(maxlen=100) keeps the last 100 log lines or events and discards older ones automatically.

We once hit a slowdown in a nightly import that matched new customer rows against existing ones. The job stored existing emails in a list and ran if email not in existing for each new row. It took about forty minutes at a few hundred thousand rows. Changing one line to build a set dropped it to seconds, with no other change.

Choosing a data structure: a decision framework

Ask these questions in order, and stop at the first yes.

  1. Do you look things up by a name or an ID? Use a dictionary. Nothing else gives you constant-time lookup by key.
  2. Do you only care whether a value is present, or need unique values? Use a set.
  3. Is it one record with a fixed number of fields? Use a tuple, or a namedtuple when names make the code clearer.
  4. Do you add and remove items at both ends? Use a collections.deque.
  5. Is it an ordered collection of similar items? Use a list.

When two answers apply, combine structures. A dictionary whose values are lists, or a list of tuples, models most real data well.

When NOT to use the built-in structures

The four built-ins cover most code, but they are the wrong choice in a few situations.

  • Large numeric arrays. A list of a million floats stores a million separate Python objects. For numeric work at that scale, a NumPy array uses far less memory and computes much faster.
  • Records with many fields and behavior. Once a tuple grows past four or five positions, or needs methods, readers cannot tell what row[6] means. Define a class instead.
  • Data that must stay sorted while it changes. Re-sorting a list after every insert is wasteful. The standard library’s bisect and heapq modules keep order or priority at lower cost.

Common mistakes

  • Testing membership in a list inside a loop. Each test scans the whole list, so the total work grows with the square of the data size. Jobs that took seconds start taking hours.
  • Changing a list while looping over it. Removing items during iteration shifts the remaining ones, so the loop skips elements. Build a new list instead.
  • Changing a dictionary’s size while looping over it. Python raises RuntimeError: dictionary changed size during iteration. Loop over list(d) if you must add or delete keys.
  • Assuming assignment copies. Writing b = a creates a second name for one object. A change through b silently corrupts a.
  • Using {} for an empty set. It creates a dictionary, and the later .add() call fails with AttributeError.
  • Using a list as a dictionary key. Lists are mutable and therefore unhashable, so Python raises TypeError. Convert the list to a tuple first.

Key takeaways

  • Pick the structure by the question you ask most: position, record, presence, or key.
  • Convert a list to a set before any loop that tests membership against it.
  • Dictionaries keep insertion order since Python 3.7, while sets have no order.
  • Assignment shares an object. Use .copy() for flat data and copy.deepcopy() for nested data.
  • Build nested lists with a comprehension, never with [[0] * n] * m.
  • Use tuples for fixed records and as dictionary keys, because they are hashable.
  • Reach for Counter, defaultdict, and deque before writing the same logic by hand.

FAQ

What are the four built-in data structures in Python?

They are the list, the tuple, the set, and the dictionary. Lists and tuples keep items in order, sets hold unique values, and dictionaries map keys to values.

What is the difference between a list and a tuple in Python?

A list is mutable and suits a collection whose length changes. A tuple is immutable and suits one fixed record. Because tuples are hashable, you can also use them as dictionary keys.

When should I use a set instead of a list?

Use a set when you need unique values or frequent membership tests. Checking whether a value exists in a set takes constant time on average, while a list must scan its items one by one.

Are Python dictionaries ordered?

Yes. Since Python 3.7, a dictionary keeps the order in which you inserted its keys. It does not sort them, so call sorted() when you need sorted output.

What is the difference between a shallow copy and a deep copy?

A shallow copy creates a new outer container that shares the inner objects. A deep copy duplicates every nested level, so changes to the copy never reach the original.

Ask the question first, then pick the container

The four built-in containers are simple, and choosing among them is a matter of knowing what you will ask of the data. A list where a set belongs still runs, which is exactly why the mistake survives until the data grows. Check your loops for hidden scans, and copy on purpose.

Rule of thumb: position means list, record means tuple, presence means set, and key means dictionary.

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *