How to Test and Evaluate LLM Applications in Python
Separate deterministic tests from evals, build a golden dataset, score with code before judges, and read pass rates with their margin of error.
A developer improves a prompt. The three examples they always try now give better answers, so the change ships. A week later, support reports that the assistant has started misclassifying refund requests, which used to work. Nobody had a list of cases that used to work. The team had been testing by looking at a few outputs and feeling good about them.
To evaluate LLM applications means to measure, with repeatable numbers, how well a model-backed feature does its job. An eval is a set of inputs with known good outcomes, a way to score each output, and an aggregate result. It differs from a unit test in one way: the answer is a rate, not a pass or a fail.
You need Python 3.13 and pytest 8. The complete example runs offline with a stand-in model, so it needs no key and costs nothing. Pointing it at a real model takes one environment variable, and from then on each run is billed per token. If pytest fixtures and parametrize are unfamiliar, read about them first.
My position: an LLM feature without an eval set is not finished. You cannot improve what you cannot measure, and with a non-deterministic component, looking at a few outputs is not measurement. However, an eval score reported without its uncertainty is nearly as misleading as no score.
Three layers: tests, evals, and monitoring
/\
/ \ PRODUCTION MONITORING
/ \ real traffic, sampled and scored; user feedback
/------\
/ \ EVALS
/ \ golden dataset + real model, a pass rate, on demand or nightly
/------------\
/ \ UNIT TESTS
/ \ fake model, deterministic, every commit, free
/------------------\
| Layer | Model | Result | Runs | Catches |
|---|---|---|---|---|
| Unit tests | A fake you control | Pass or fail | Every commit | Bugs in your parsing, prompts, retries, and control flow |
| Evals | The real model | A pass rate against a threshold | Before merging prompt or model changes, and nightly | Quality regressions |
| Monitoring | The real model on real traffic | Trends over time | Continuously | Inputs you never imagined |
Most teams have a little of the first layer and none of the second. The second layer is where quality lives.
The myth: you cannot test something non-deterministic
People often say that LLM features are untestable because the output changes from run to run. The claim confuses two things. Your code around the model is fully deterministic: it builds a prompt, parses a reply, validates fields, and decides what to do. You can test all of that exactly, by replacing the model with a fake.
The model’s behavior is statistical, and statistics is the right tool for it. You do not ask whether one output is identical to a stored string. You ask how many of 200 cases met the requirement, and whether that share went up or down.
A complete example to evaluate LLM applications
The feature under test classifies a support message into one of four categories. Save this file as evals.py.
import math
import os
import sys
from collections.abc import Callable
from dataclasses import dataclass
MODEL = 'gpt-4o-mini'
THRESHOLD = 0.8
CATEGORIES = ('billing', 'bug', 'account', 'other')
PROMPT = (
'Classify the support message into exactly one category: '
'billing, bug, account, or other. Reply with the category word only.'
)
Generate = Callable[[str, str], str] # (system prompt, user message) -> reply text
# --- the application code under test ---------------------------------------
def classify(generate: Generate, message: str) -> str:
raw = generate(PROMPT, message)
answer = raw.strip().lower().strip('.')
return answer if answer in CATEGORIES else 'other'
# --- the golden dataset ------------------------------------------------------
@dataclass(frozen=True)
class Case:
id: str
message: str
expected: str
DATASET = [
Case('double-charge', 'I was charged twice for my plan this month.', 'billing'),
Case('crash-on-save', 'The app crashes every time I press save.', 'bug'),
Case('reset-password', 'I cannot log in and need to reset my password.', 'account'),
Case('refund-request', 'Please refund my last payment.', 'billing'),
Case('feature-idea', 'It would be nice to have a dark theme.', 'other'),
Case('wrong-total', 'The invoice total on the page is calculated wrong.', 'bug'),
]
# --- models ------------------------------------------------------------------
def keyword_model(system: str, user: str) -> str:
"""Offline stand-in. Deterministic, free, and deliberately imperfect."""
text = user.lower()
if any(word in text for word in ('charged', 'refund', 'invoice', 'payment')):
return 'Billing.'
if any(word in text for word in ('crash', 'error', 'wrong')):
return 'bug'
if any(word in text for word in ('password', 'log in', 'account')):
return 'account'
return 'other'
def make_openai_model() -> Generate:
"""Real model. Needs OPENAI_API_KEY, and every case is a billed call."""
from openai import OpenAI
client = OpenAI(timeout=30.0, max_retries=2)
def generate(system: str, user: str) -> str:
response = client.chat.completions.create(
model=MODEL,
messages=[
{'role': 'system', 'content': system},
{'role': 'user', 'content': user},
],
temperature=0,
max_tokens=5,
)
return response.choices[0].message.content or ''
return generate
# --- the eval runner ---------------------------------------------------------
@dataclass(frozen=True)
class Result:
case: Case
actual: str
passed: bool
def run_eval(generate: Generate, dataset: list[Case]) -> list[Result]:
results = []
for case in dataset:
actual = classify(generate, case.message)
results.append(Result(case, actual, actual == case.expected))
return results
def pass_rate(results: list[Result]) -> float:
return sum(result.passed for result in results) / len(results)
def margin_of_error(rate: float, total: int) -> float:
"""Approximate 95 percent margin for a pass rate measured on `total` cases."""
return 1.96 * math.sqrt(rate * (1 - rate) / total)
def main() -> int:
use_real = os.environ.get('EVAL_PROVIDER') == 'openai'
generate = make_openai_model() if use_real else keyword_model
results = run_eval(generate, DATASET)
for result in results:
if result.passed:
print(f'PASS {result.case.id}')
else:
print(
f'FAIL {result.case.id}: '
f'expected {result.case.expected!r}, got {result.actual!r}'
)
rate = pass_rate(results)
passed = sum(result.passed for result in results)
margin = margin_of_error(rate, len(results))
print(f'pass rate: {rate:.2f} ({passed} of {len(results)}), 95% margin: +/- {margin:.2f}')
return 0 if rate >= THRESHOLD else 1
if __name__ == '__main__':
raise SystemExit(main())
python evals.py
PASS double-charge
PASS crash-on-save
PASS reset-password
PASS refund-request
PASS feature-idea
FAIL wrong-total: expected 'bug', got 'billing'
pass rate: 0.83 (5 of 6), 95% margin: +/- 0.30
The offline model gets five of six right. The script exits with status 0, because 0.83 clears the threshold of 0.8. In CI, a result below the threshold would exit with status 1 and fail the build.
To run the same dataset against a real model, install the SDK, set your key in an environment variable, and select the provider. Each case is a paid API call, and results can vary slightly between runs.
# macOS and Linux
python -m venv .venv && source .venv/bin/activate
python -m pip install openai pytest
export OPENAI_API_KEY="sk-your-key-here"
EVAL_PROVIDER=openai python evals.py
On Windows PowerShell, activate with .venv\Scripts\Activate.ps1, and set $env:OPENAI_API_KEY and $env:EVAL_PROVIDER. The model name is one constant at the top, since model names change quickly.
What the pieces are
- The application function takes its model as an argument.
classifydoes not know whether it is talking to a real model or a stand-in. That single design choice makes everything else possible. - The dataset is data. Each case has an ID, an input, and an expected outcome. IDs matter: “wrong-total failed” is a bug report, and “case 5 failed” is a chore.
- The scorer is a comparison. Here it is equality, the simplest possible scorer.
- The report shows failures first-class. You read the failing cases, not only the rate.
Look at the failing case. “The invoice total on the page is calculated wrong” mentions an invoice and describes a defect. Is it billing or a bug? Reasonable people disagree. Cases like this are valuable. They force you to write down the rule, and writing down the rule usually improves the prompt.
Read every score with its margin of error
Here is the insight that most eval write-ups leave out. A pass rate is an estimate from a sample, and small samples are noisy. The last line of the output says so directly: 0.83, plus or minus 0.30. With six cases, the true rate could plausibly be anywhere from about 0.5 to 1.0. That number supports no decision at all.
The margin shrinks with the square root of the number of cases. For a measured pass rate of 90 percent:
| Cases | Approximate 95 percent margin | What you can conclude |
|---|---|---|
| 25 | Plus or minus 12 points | Only that the feature is not badly broken |
| 50 | Plus or minus 8 points | Large regressions |
| 100 | Plus or minus 6 points | Differences of 10 points or more |
| 400 | Plus or minus 3 points | Differences of about 5 points |
| 1,000 | Plus or minus 2 points | Differences of about 3 points |
These values come from the margin_of_error function in the script, which uses the normal approximation. It is rough for very small samples, and it is good enough to stop the most common mistake: declaring that a new prompt is better because it scored 92 percent against 88 percent on 25 cases. That gap is well inside the noise. A second run of the same prompt could reverse it.
Two practical rules follow. First, when you compare two versions, look at which cases changed, not only at the two totals. Five cases newly fixed and five newly broken is a very different story from zero and zero. Second, grow the dataset until the margin is smaller than the difference you care about.
Build a golden dataset that earns trust
- Start with real inputs. Pull messages from logs, tickets, or search queries, with personal data removed. Invented examples are too clean.
- Cover the edges on purpose. Include empty input, very long input, other languages, ambiguous cases, and attempts at prompt injection.
- Every bug becomes a case. When production produces a wrong answer, add that input and its correct outcome before you fix anything. The dataset then grows exactly where the feature is weak.
- Have a person write the expected answers. If a model writes them, you are measuring agreement between models, not correctness.
- Keep a held-out portion. If you tune the prompt against every case, you overfit to them. Reserve some cases that you look at only before a release.
- Version it with the code. The dataset lives in the repository, so that a score always refers to a known set of cases.
Thirty well-chosen cases are enough to begin. They will not give a tight margin, yet they will catch the embarrassing regressions on day one.
Choose the cheapest scorer that works
| Scorer | Checks | Cost | Use it for |
|---|---|---|---|
| Exact match | Output equals the expected value | Free | Classification, routing, extraction of fixed values |
| Contains or regular expression | Required facts or forbidden phrases appear | Free | Answers that must mention a number, a name, or a refusal sentence |
| Schema validation | Output parses into the expected structure | Free | Structured output |
| Grounding check | Quoted values exist in the source | Free | Extraction and retrieval answers |
| State check | The system ended in the right state | Free | Agents and tool use |
| Model as judge | A second model applies a rubric | One more model call per case | Tone, helpfulness, and faithfulness of free text |
| Human review | A person applies the rubric | Slow and expensive | Calibrating everything above |
Work down the table, and stop at the first row that answers your question. Code-based scorers are free, instant, and perfectly consistent. For many features, a surprising amount of quality can be checked that way.
Different features have natural metrics. For retrieval, measure whether the right document came back before you measure the answer, as in the retrieval hit rate in a RAG pipeline. For an agent loop, check the final state of the system and the number of steps taken, not the wording of the final message.
Using a model as a judge, carefully
Some qualities resist code. Is this summary faithful to the document? Is this reply polite? For those, you can ask a second model to grade the output against a rubric. This is called LLM-as-judge. It is useful, and it is not objective.
import json
from collections.abc import Callable
JUDGE_PROMPT = (
'You are grading an answer against a reference. '
'Reply with JSON only: {"reason": "one sentence", "verdict": "pass" or "fail"}. '
'Pass only if the answer states the same facts as the reference and adds no '
'unsupported claims. Ignore differences in wording and length.'
)
def judge(generate: Callable[[str, str], str], question: str, answer: str, reference: str) -> bool:
user = f'Question: {question}\nReference: {reference}\nAnswer: {answer}'
raw = generate(JUDGE_PROMPT, user)
try:
verdict = json.loads(raw).get('verdict')
except (json.JSONDecodeError, AttributeError):
return False # an unreadable verdict counts as a failure
return verdict == 'pass'
def test_judge_parses_a_pass():
fake = lambda system, user: '{"reason": "Same facts.", "verdict": "pass"}'
assert judge(fake, 'How long?', 'Five business days.', 'Refunds take five business days.')
def test_unreadable_verdict_counts_as_fail():
fake = lambda system, user: 'Looks good to me!'
assert not judge(fake, 'How long?', 'Soon.', 'Refunds take five business days.')
The two tests show that the judging code itself is ordinary, testable Python. The rubric asks for the reason before the verdict, so that the model commits to its reasoning first. It asks for a binary verdict, because pass or fail is far more consistent than a score from 1 to 10.
Judges have known biases, and you should design around them.
- Verbosity bias. Longer answers tend to be rated higher, even when the extra text adds nothing.
- Position bias. When comparing two answers, a judge often favors whichever comes first. Run comparisons in both orders.
- Self-preference. A model tends to rate text in its own style favorably. Using a different model as the judge reduces this.
- Inconsistency. The same input can receive different verdicts on different runs.
- Leniency. Judges pass plausible answers too easily unless the rubric lists concrete failure conditions.
The rule that keeps a judge honest is calibration. Have a person label fifty to a hundred outputs as pass or fail, run the judge on the same outputs, and measure how often they agree. If agreement is high, you can trust the judge at scale. If it is not, fix the rubric before you trust any number it produces. A judge is a model-backed feature too, so it needs its own eval.
Stronger models, including the newer reasoning models, often make more careful judges, at a higher cost per case. Keep the judge’s model name in a constant, and record it with every score. A score from one judge model is not comparable with a score from another.
Wire it into pytest and CI
Unit tests and evals belong in the same test suite, with the expensive part switched off by default. Save this as test_evals.py.
import os
import pytest
from evals import DATASET, THRESHOLD, classify, keyword_model, pass_rate, run_eval
def test_reply_is_normalized():
assert classify(lambda system, user: ' Billing.\n', 'anything') == 'billing'
def test_unexpected_reply_falls_back_to_other():
assert classify(lambda system, user: 'I think this is about money', 'anything') == 'other'
def test_offline_baseline_meets_threshold():
assert pass_rate(run_eval(keyword_model, DATASET)) >= THRESHOLD
@pytest.mark.skipif(
os.environ.get('RUN_LLM_EVALS') != '1',
reason='set RUN_LLM_EVALS=1 to run evals against the real model (billed)',
)
def test_real_model_meets_threshold():
from evals import make_openai_model
assert pass_rate(run_eval(make_openai_model(), DATASET)) >= THRESHOLD
python -m pytest -q test_evals.py
...s [100%]
3 passed, 1 skipped in 0.04s
The first two tests are true unit tests: they check parsing with tiny fakes. The third is a smoke test of the eval machinery. The fourth calls the real model, and it runs only when you ask for it with RUN_LLM_EVALS=1.
A workable CI policy looks like this. Unit tests run on every commit. The real-model eval runs when a pull request touches a prompt, a model name, or retrieval code, and again on a nightly schedule. The nightly run matters, because a hosted model can change behind the same name, and your code did nothing to cause the regression.
Set the threshold slightly below your current measured rate, not at an aspirational 100 percent. A threshold the suite can never meet gets disabled within a week.
Eval tools, and when to adopt one
| Tool | Style | Strength |
|---|---|---|
| Plain pytest, as above | Your own Python | No dependency, full control, easy to understand |
| promptfoo | Configuration files and a command-line tool | Side-by-side comparison of prompts and models, with a results viewer |
| DeepEval | pytest-style Python library | Ready-made metrics, including judge-based ones |
| Ragas | Python library | Metrics for retrieval pipelines, such as faithfulness and context precision |
| Inspect | Python framework | Structured evaluations with datasets, solvers, and scorers, built for rigorous model testing |
Start with the plain version. Sixty lines of your own code teach you what a dataset, a scorer, and a threshold are. Adopt a tool when you need what it adds: a comparison interface, a library of metrics, or shared reports. The concepts transfer unchanged.
How real systems evaluate LLM features
- A dataset per feature, in the repository. Classification, extraction, and answer generation each have their own cases and their own threshold.
- Results stored per run. Each eval run records the prompt version, the model name, the pass rate, and the per-case results, so that any two runs can be compared.
- Case-level diffs in code review. A pull request that changes a prompt shows which cases flipped from pass to fail, and the reverse.
- Sampled production scoring. A small share of live requests is scored by code checks or a calibrated judge, and the trend is charted.
- User signals feed the dataset. Thumbs-down ratings and support escalations are reviewed weekly, and the confirmed failures become new cases.
In my experience reviewing a prompt change that “clearly improved” a summarization feature, the author had checked four favorite examples. We ran the existing set of 120 cases, and the overall rate barely moved, from 86 to 87 percent. The per-case diff told the real story: eleven cases improved and ten got worse, most of them long documents that the new, shorter instructions handled badly. We kept the new wording for short inputs only. Without the per-case view, we would have shipped a silent regression.
Deciding what to measure: a decision framework
- Is the thing you are checking your own code? Write a unit test with a fake model. It should never call an API.
- Does the output have one right answer? Use exact match or a normalizing comparison.
- Must the output contain or avoid specific content? Use contains checks or regular expressions.
- Is the output structured? Validate the schema, then check values against the source.
- Is quality a matter of judgment? Use a model as judge with a binary rubric, and calibrate it against human labels first.
- Is the difference you care about smaller than the margin? Add cases before you draw a conclusion.
When NOT to rely on automated evals
- High-stakes outputs. For medical, legal, or financial content, automated scores are a filter. A qualified person still reviews samples.
- A feature that is still changing shape. If the task definition changes weekly, a large dataset goes stale as fast as you build it. Keep a small one until the requirements settle.
- Qualities the judge has not been calibrated on. An uncalibrated judge produces confident numbers that mean nothing. No number is better than a wrong one.
Common mistakes
- Testing by eyeballing a few outputs. Favorite examples improve while everything else regresses unseen.
- Asserting exact strings from a live model. The test fails at random, the team marks it flaky, and it gets deleted.
- Calling the real API in unit tests. The suite becomes slow, costly, and dependent on a provider’s uptime.
- Comparing scores from tiny datasets. A four-point difference on 25 cases is noise, and decisions made on it are coin flips.
- Tuning on the whole dataset. The prompt overfits to the cases, the score rises, and real users see no improvement.
- Trusting an uncalibrated judge. The judge rewards length and fluency, and a worse system scores higher.
Key takeaways
- Unit-test your own code with fake models, and evaluate model behavior with a golden dataset.
- Inject the model into your functions, so that both kinds of test are possible.
- An eval result is a rate with a margin of error, and small datasets give wide margins.
- Compare versions case by case, not only by their totals.
- Prefer code-based scorers, and use a model as judge only for subjective qualities.
- Calibrate any judge against human labels before trusting it.
- Run evals on prompt and model changes and on a schedule, and turn every production failure into a case.
FAQ
How do you test an LLM application?
Test your own code with unit tests that replace the model with a fake. Evaluate the model’s behavior separately by running a dataset of inputs with known outcomes through the real model and measuring the pass rate.
What is an LLM eval?
An eval is a repeatable measurement of quality. It consists of a set of test inputs, a method for scoring each output, and an aggregate result such as a pass rate that you compare with a threshold.
What is LLM-as-a-judge?
It is the practice of using a second language model to grade outputs against a rubric. It suits subjective qualities such as faithfulness and tone, and it must be calibrated against human judgments, because judges have biases.
How many test cases do I need for an LLM eval?
Thirty cases catch obvious regressions. To detect a difference of about five percentage points reliably, you need several hundred. The margin of error shrinks with the square root of the number of cases.
How do I test non-deterministic LLM output?
Do not assert exact strings. Check properties of the output, such as a category, a required fact, or a valid structure, across many cases, and require that the pass rate stays above a threshold.
Measure the rate, read the failures
Testing a model-backed feature is not harder than testing other software. It is different in one respect: part of the answer is a percentage. Keep the deterministic part under ordinary unit tests, give the statistical part a dataset and a threshold, and look at the individual failures every time the number moves.
Rule of thumb: if you cannot say which cases a change fixed and which it broke, you have not evaluated it.
