How Do LLMs Work? A Developer’s Guide to Tokens, Context, and Embeddings

An LLM predicts the next token, over and over. Learn tokens, context windows, temperature, and embeddings through small Python programs you can run locally.

Executive Summary: An LLM predicts the next token from a sequence of tokens and repeats until the answer is done. This guide explains what follows for developers: tokens drive cost and limits, the context window is the model’s only memory, temperature controls randomness, and embeddings make text comparable, and why your code must still check what the model returns.

Ask a chat model to multiply two six-digit numbers, and it often answers instantly, confidently, and wrongly. Ask it to write a function that multiplies them, and the code is usually correct. That odd pair of results makes sense once you know what the model is doing. It is not calculating. It is predicting which text is likely to come next.

So, how do LLMs work? A large language model (LLM) is a neural network trained on a very large amount of text to predict the next token in a sequence. Chat assistants such as ChatGPT, built on GPT-3.5 and GPT-4, and Claude are applications wrapped around that one ability.

This guide is for developers who want an accurate mental model before they build with these systems. It uses Python 3.11, and you need no API key and no account. Two examples use only the standard library, and one uses the open source tiktoken package. Python is the main language of this field, which is one reason to learn where Python is used.

My position: you do not need the mathematics of neural networks to build reliable LLM features. You do need four concepts, tokens, context, sampling, and embeddings, and you need them precisely. Most early mistakes come from a wrong picture of one of the four.

How do LLMs work: the core loop

Strip away the chat interface, and one loop remains.

"The cat sat on the"
        |
        v
1. Tokenize         text -> pieces -> integer IDs
        |
        v
2. Model            reads ALL tokens so far, outputs a score for every
                    token in its vocabulary (tens of thousands of scores)
        |
        v
3. Sample           turn scores into probabilities, pick one token
                    " mat" 65%   " sofa" 24%   " roof" 9%   ...
        |
        v
4. Append           "The cat sat on the mat"
        |
        +------> back to step 2, until a stop token or a length limit

The model produces one token per pass. A 500-word answer is roughly 700 passes through the network, each one reading everything written so far. Consequently, long outputs are slow, and you see chat interfaces print text piece by piece.

The network inside step 2 is a transformer, an architecture introduced in the 2017 paper “Attention Is All You Need”. Its key mechanism, attention, lets each token’s representation draw on every other token in the input. You can use an LLM well without knowing more than that.

Two consequences deserve attention now. First, the model has no separate store of facts that it looks up. Knowledge is spread through billions of numeric weights fixed at training time. Second, nothing in the loop checks whether the output is true. The model selects likely text.

Tokens: the unit of everything

Models do not read characters or words. They read tokens, which are chunks of text from a fixed vocabulary. Common words are one token, and rarer words split into several pieces. For English text, a token averages about four characters, or about three quarters of a word.

You can inspect real tokenization locally with tiktoken, the open source tokenizer library from OpenAI. It needs no API key. It downloads a small vocabulary file the first time you use an encoding.

python --version
python -m venv .venv
source .venv/bin/activate          # macOS and Linux
.venv\Scripts\Activate.ps1         # Windows PowerShell
python -m pip install tiktoken

Save this as tokens.py.

import sys

import tiktoken

MODEL = 'gpt-3.5-turbo'


def get_encoding(model: str) -> tiktoken.Encoding:
    try:
        return tiktoken.encoding_for_model(model)
    except KeyError:
        return tiktoken.get_encoding('cl100k_base')


def main() -> None:
    text = ' '.join(sys.argv[1:]) or 'hello world'
    encoding = get_encoding(MODEL)
    token_ids = encoding.encode(text)
    pieces = [encoding.decode([token_id]) for token_id in token_ids]
    print(f'characters: {len(text)}')
    print(f'tokens:     {len(token_ids)}')
    print(f'ids:        {token_ids}')
    print(f'pieces:     {pieces}')


if __name__ == '__main__':
    main()
python tokens.py
characters: 11
tokens:     2
ids:        [15339, 1917]
pieces:     ['hello', ' world']

Notice that the space belongs to the second token. Now pass your own text as arguments, such as a long technical word, a line of code, or a sentence in another language. You will see three patterns:

  • Rare and long words split into several tokens.
  • Code and JSON use more tokens per character than prose, because of punctuation and unusual names.
  • Many non-English languages use more tokens for the same meaning.

The try block handles a model name the library does not recognize, which raises KeyError. Model names change quickly, so keep the name in one constant, as this script does.

Why tokens matter to your design

Tokens are the unit for three things at once. Providers bill by tokens, both for input and for output. Limits are stated in tokens. Latency grows with the number of output tokens, since each one is a separate pass. Therefore, count tokens in code rather than guessing from characters, and count before you send a request.

Tokenization also explains the arithmetic failure from the introduction. A long number is split into arbitrary multi-digit pieces, so the model never sees digits lined up in columns. It predicts plausible-looking digits instead.

The context window: the model’s only memory

The context window is the maximum number of tokens a model can handle in one request, counting the input and the output together. At the time of writing, published limits look like this:

Model Context window Rough size in English words
GPT-3.5 Turbo 4,096 tokens About 3,000
GPT-4 8,192 tokens, with a 32,768 token variant in limited access About 6,000, or 24,000
Claude (100K context, announced this month) About 100,000 tokens About 75,000

These numbers will change, so check each provider’s documentation when you build.

The myth: the model remembers your conversation

A chat interface feels like talking to something with memory. The model has none. Each request is independent, and the model keeps nothing between calls. A chat application creates the illusion by sending the whole conversation again with every new message.

Request 1:  [system prompt] [user: "My name is Asha"]
Request 2:  [system prompt] [user: "My name is Asha"] [assistant: "Hi Asha"]
            [user: "What is my name?"]

The model answers the second request correctly only because the first exchange is included in it. Three practical facts follow:

  • Every turn costs more than the last, because the history grows.
  • When the history exceeds the window, something must be dropped or summarized, and the model “forgets” it.
  • Anything you want the model to know, such as a document or today’s date, must be in the request.

The text you send is called the prompt. In chat models, it is a list of messages with roles: a system message that sets behavior, user messages, and earlier assistant replies. To the model, it all becomes one long token sequence.

Budget the window in code

Here is the rule I apply to every LLM feature, and the insight most first projects lack. The window is a budget with three parts, and you must reserve the output’s share before you fill the input:

fixed prompt tokens + variable content tokens + reserved output tokens  <=  context window

If you fill the window with input, the model has no room to answer, and the reply is cut off mid-sentence. The script below trims a conversation to fit a budget. It uses only the standard library, and it takes the token counter as a parameter, so its test needs no network and no tokenizer. Save it as budget.py.

import unittest
from collections.abc import Callable

CONTEXT_WINDOW = 4096
RESERVED_FOR_OUTPUT = 500


def rough_count(text: str) -> int:
    """Estimate tokens as one per four characters. Use a real tokenizer in production."""
    return max(1, len(text) // 4)


def fit_history(
    system_prompt: str,
    history: list[str],
    count: Callable[[str], int] = rough_count,
    window: int = CONTEXT_WINDOW,
    reserved: int = RESERVED_FOR_OUTPUT,
) -> list[str]:
    """Keep the most recent messages that fit, always keeping the system prompt."""
    budget = window - reserved - count(system_prompt)
    if budget <= 0:
        raise ValueError('system prompt and reserved output exceed the context window')
    kept: list[str] = []
    for message in reversed(history):
        cost = count(message)
        if cost > budget:
            break
        kept.append(message)
        budget -= cost
    return list(reversed(kept))


class FitHistoryTest(unittest.TestCase):
    def test_drops_oldest_messages_first(self) -> None:
        one_token_per_word = lambda text: len(text.split())
        history = ['first message here', 'second message here', 'third message here']
        kept = fit_history('be brief', history, one_token_per_word, window=12, reserved=4)
        self.assertEqual(kept, ['second message here', 'third message here'])

    def test_rejects_an_impossible_budget(self) -> None:
        with self.assertRaises(ValueError):
            fit_history('x' * 100, [], window=10, reserved=10)


if __name__ == '__main__':
    unittest.main()
python budget.py
..
----------------------------------------------------------------------
Ran 2 tests in 0.000s

OK

In the first test, the window is 12 tokens, 4 are reserved for output, and the system prompt uses 2. That leaves 6, which fits the two newest three-word messages and drops the oldest. Dropping old messages is the simplest strategy. Real applications also summarize older turns, or fetch only the relevant passages.

Temperature: why the same prompt gives different answers

The model outputs a score for every token in its vocabulary. A function called softmax turns the scores into probabilities. The application then picks a token at random, weighted by those probabilities. That random pick is called sampling, and it is why model output is non-deterministic: the same prompt can produce different text each time.

Temperature is a number that reshapes the probabilities before sampling. The scores are divided by the temperature. A low value sharpens the distribution toward the top choice, and a high value flattens it. This standard library script shows the effect on four made-up candidate tokens. Save it as temperature.py.

import math
import random

CANDIDATES = ['mat', 'sofa', 'roof', 'moon']
SCORES = [4.0, 3.0, 2.0, 0.5]


def softmax(scores: list[float], temperature: float) -> list[float]:
    if temperature <= 0:
        raise ValueError('temperature must be greater than zero')
    scaled = [score / temperature for score in scores]
    highest = max(scaled)
    weights = [math.exp(value - highest) for value in scaled]
    total = sum(weights)
    return [weight / total for weight in weights]


def main() -> None:
    for temperature in (0.2, 1.0, 2.0):
        probabilities = softmax(SCORES, temperature)
        summary = ', '.join(
            f'{word} {probability:.2f}'
            for word, probability in zip(CANDIDATES, probabilities)
        )
        print(f'temperature {temperature}: {summary}')
    picks = random.choices(CANDIDATES, weights=softmax(SCORES, 1.0), k=10)
    print('ten samples at 1.0:', ' '.join(picks))


if __name__ == '__main__':
    main()
python temperature.py
temperature 0.2: mat 0.99, sofa 0.01, roof 0.00, moon 0.00
temperature 1.0: mat 0.65, sofa 0.24, roof 0.09, moon 0.02
temperature 2.0: mat 0.47, sofa 0.28, roof 0.17, moon 0.08
ten samples at 1.0: mat mat sofa mat roof mat mat sofa mat mat

The first three lines are always the same. The last line changes on every run, which is the point. At 0.2, the model almost always says “mat”. At 2.0, “moon” appears about one time in twelve.

Temperature Behavior Use it for
0 to 0.3 Focused and repeatable, though not guaranteed identical Extraction, classification, code, anything you parse
0.7 to 1.0 Varied and natural Drafting and conversation
Above 1.0 Increasingly erratic Brainstorming, with a human reading the result

Low temperature does not make a model truthful. It makes the model pick its most likely token more consistently, and the most likely token can be wrong.

Embeddings: text as coordinates

An embedding is a list of numbers, a vector, that represents the meaning of a piece of text. An embedding model maps text to a point in a space with hundreds or thousands of dimensions. Texts with similar meanings land near each other, even when they share no words.

You compare two embeddings with cosine similarity, which measures the angle between the vectors. A value near 1 means similar direction, and a value near 0 means unrelated. The script below uses tiny hand-made vectors, stored in ordinary lists and dictionaries, to show the calculation. Real embeddings come from a model and have far more dimensions, for example 1,536.

import math

VECTORS = {
    'refund policy': [0.9, 0.1, 0.0],
    'money back guarantee': [0.8, 0.2, 0.1],
    'server uptime report': [0.1, 0.0, 0.9],
}


def cosine_similarity(a: list[float], b: list[float]) -> float:
    dot = sum(x * y for x, y in zip(a, b))
    norm_a = math.sqrt(sum(x * x for x in a))
    norm_b = math.sqrt(sum(y * y for y in b))
    return dot / (norm_a * norm_b)


def main() -> None:
    query = 'refund policy'
    for text, vector in VECTORS.items():
        if text != query:
            score = cosine_similarity(VECTORS[query], vector)
            print(f'{text}: {score:.2f}')


if __name__ == '__main__':
    main()
money back guarantee: 0.98
server uptime report: 0.11

“Refund policy” and “money back guarantee” share no words, and they score 0.98. A keyword search would miss that match. This is the basis of semantic search: embed your documents once, embed the user’s question, and return the nearest documents. It is also how applications find the right passages to place in a limited context window.

Embeddings and chat models are different tools. An embedding model returns a vector and generates no text. A chat model generates text and returns no vector. Many systems use both.

Hallucination: fluent is not the same as true

A hallucination is output that reads well and is false: an invented citation, a function that does not exist, or a confident wrong date. It is not a malfunction. The loop selects likely tokens, and a plausible false sentence can be more likely than “I do not know”.

Hallucination becomes more probable in predictable situations:

  • The question concerns events after the model’s training data ends.
  • The question concerns private information, such as your codebase or your company’s policies.
  • The request needs exact recall, such as URLs, quotations, version numbers, and legal references.
  • The request needs exact computation.

You reduce it by supplying the facts in the prompt, asking the model to answer only from that material, and checking the output in code. You do not eliminate it.

How real systems use these four ideas

  • Retrieval before generation. Applications embed their documents, find the passages nearest to the question, and paste those passages into the prompt. The model answers from supplied text instead of from memory.
  • Token accounting on every call. Services count tokens before sending, log input and output counts per request, and alert on unusual totals. The token count is the bill.
  • History management. Chat products trim or summarize older turns to stay inside the window, and they pin the system prompt so that it is never dropped.
  • Low temperature for machine-read output. When code parses the reply, teams use a low temperature and still validate the result, because repeatable is not the same as correct.
  • Output limits and timeouts. Every request sets a maximum output length. Since each token takes time to generate, the limit bounds both cost and latency.

We once hit a bug when a support chatbot prototype worked in every demo and then failed for real users after about fifteen messages. The API returned an error saying the model’s maximum context length was 4097 tokens and our messages had exceeded it. Nobody had counted tokens, because short test conversations never came close. We added a budget function like the one above and a test with a long conversation, and the failure never returned.

Deciding whether an LLM fits your feature: a decision framework

  1. Is the task about language? Summarizing, rewriting, classifying, extracting, and drafting are good fits. Arithmetic, sorting, and exact lookups are not.
  2. Can the answer be checked? If code or a person can verify the output cheaply, proceed. If a wrong answer would go unnoticed and cause harm, add a review step or stop.
  3. Does it need facts the model cannot know? Plan retrieval: find the relevant text and put it in the prompt.
  4. Does it fit the window? Estimate tokens for the prompt, the content, and the output. If it does not fit, you need chunking or a different design.
  5. Can you afford the variability? Output is non-deterministic. If you need the same result every time, use ordinary code.

When NOT to use an LLM

  • Exact computation and deterministic rules. Totals, tax calculations, date arithmetic, and validation rules belong in code. Code is faster, cheaper, and right every time.
  • Lookups of known data. If the answer is in your database, query the database. A model may return a plausible value that is not the stored one.
  • High-stakes decisions without review. Medical, legal, and financial conclusions need a qualified person. A fluent paragraph is not evidence.

Common mistakes

  • Assuming the model remembers earlier requests. Each call is independent. Leave out the history, and the model has no idea what “it” refers to.
  • Estimating size in characters or words. Limits and bills are in tokens. Code and non-English text can use far more tokens than you expect.
  • Filling the window with input. No room remains for the answer, and replies are truncated mid-sentence.
  • Trusting fluent output. Confident wording says nothing about accuracy. Unverified answers ship invented facts to users.
  • Expecting identical output at low temperature. Low temperature reduces variation. It does not guarantee the same text, so tests that compare exact strings fail at random.
  • Hard-coding a model name everywhere. Models are replaced and renamed often. Keep the name and its window size in one place.

Key takeaways

  • An LLM predicts the next token and repeats. Everything else is built on that loop.
  • Tokens, not words, determine cost, limits, and speed. Count them in code.
  • The context window is the model’s only memory, and it covers input plus output.
  • The model is stateless. Your application resends whatever it should know.
  • Temperature controls randomness, not truthfulness.
  • Embeddings turn text into vectors, and cosine similarity finds related text.
  • Fluent output can be false, so supply facts in the prompt and verify results.

FAQ

How does a large language model generate text?

It converts the input into tokens, computes a probability for every possible next token, picks one, appends it, and repeats. The answer is produced one token at a time.

What is a token in an LLM?

A token is a chunk of text from the model’s vocabulary, often a word or part of a word. In English, one token is about four characters on average. Models measure limits and usage in tokens.

What is a context window?

The context window is the maximum number of tokens a model can process in one request, including both the prompt and the generated reply. Text outside the window is invisible to the model.

What does temperature do in an LLM?

Temperature adjusts how random the choice of each next token is. Low values make output focused and repeatable, and high values make it more varied.

Why do LLMs hallucinate?

The model selects likely text and has no step that checks facts. When it lacks the needed information, a plausible but false continuation can still be the most probable one.

Predict, do not assume: build on what the model actually does

An LLM is easier to work with once it stops being mysterious. It predicts tokens from the tokens you give it, within a fixed window, with controlled randomness. Design around those facts: count tokens, supply the knowledge, and check the answers.

Rule of thumb: if it is not in the context window, the model does not know it, and if your code did not check it, you do not know it either.

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *