Writing Good Tests: Patterns, Fixtures, and Anti-Patterns
Executive Summary
Writing good tests starts with the test pyramid, which sets the ratios. It calls for many fast unit tests for logic, fewer integration tests for seams, and a thin end-to-end band for critical paths. Each tier up is slower and flakier. So inverted pyramids, heavy end-to-end and light unit, produce suites nobody runs. Naming is documentation. Prefer method names that state behavior, rejectsOrderWhenPaymentDeclines, over test1. Use @Nested classes to group the facets of one unit. Inside the method, the given, when, then shape makes the three phases visible at a glance. Fixtures are the cure for hard setup. They include @BeforeEach building known state, test builders with fluent defaults for wide constructors, and a small honest fake repository shared across tests. However, use the Object Mother pattern carefully, because a shared factory grows into a change-amplifier.
Flakiness is the suite killer, and its standard causes have standard cures. Avoid Thread.sleep, and inject a Clock or an executor so time and scheduling are deterministic. Also, seed randomness, and let JUnit’s per-test lifecycle kill shared state. The practical pyramid guidance and the xUnit patterns catalog cover the rest of the vocabulary. Finally, one rule decides everything: fix or delete an untrusted test the week it loses trust.
The Test Pyramid: Ratios That Keep Suites Fast
The pyramid is a ratio statement. The ratios exist because each tier up costs more of the two currencies that matter, time and trust:
/ \ few end-to-end tests
/----\ (whole system, real dependencies,
/ \ the slowest, flakiest tier)
/--------\
/ \ fewer integration tests
/--------------\ (seams: database, HTTP, real fakes)
/ \
/--------------------\ many unit tests
/______________________\ (logic only: milliseconds, deterministic)
| Tier | Runs against | Speed | Counts as |
|---|---|---|---|
| Unit | Logic, with mocks or fakes at the boundary | Milliseconds, thousands per second | The bulk of the suite |
| Integration | Real seams: a database, the HTTP client, message broker | Seconds each | A focused band, tagged and named |
| End-to-end | The deployed shape of the system | Minutes, environment-dependent | Critical paths only |
The inversion is the failure mode to recognize. A team that writes mostly end-to-end tests produces a suite that takes an hour and flakes on environment weather. As a result, the suite runs only on CI, where the team dismisses a red build as “probably flaky” by Thursday. The pyramid is not dogma. Rather, it is arithmetic: push each verification to the cheapest tier that can honestly perform it. Then the suite stays fast enough to be trusted.
Writing Good Tests: Naming and Structure as Documentation
Writing good tests means building a suite that answers “what does this class do” without reading it. The naming carries that weight:
@Nested class WhenPaymentDeclines { // the facet under one group
@Test
void refundsNothingAndMarksOrderDeclined() { ... } // the behavior, as a sentence
}
@Nested class WhenGatewayTimesOut {
@Test
void surfacesGatewayTimeoutToCaller() { ... } // the exceptions article's path
}
Inside the method, the given, when, then shape keeps the phases honest. Given an order over the limit, when the service places it, then the result is declined and no charge occurred. The shape matters because tests fail, and the reader of a failure needs to find the phase that broke in seconds, not minutes. Two rules complete the discipline. First, test one behavior per test, so a failure names one problem. Second, never mention implementation in names, like calls repository then updates cache. Implementation details change, and behavior should not.
Fixtures and Builders: Taming Hard Setup
A fixture is known state a test runs against. The moment construction gets wide, with eight constructor arguments and three nulls, the suite drowns in setup. The builder pattern fixes it with fluent defaults:
// the test builder: defaults that are valid, overridden per test
class OrderBuilder {
private String token = "tok_valid";
private BigDecimal total = new BigDecimal("19.99");
OrderBuilder withToken(String token) { this.token = token; return this; }
OrderBuilder overLimit() { this.total = new BigDecimal("9999.00"); return this; }
Order build() { return new Order(token, total); }
}
// in the test: setup is one readable line, and the interesting value is the
// only one mentioned
var order = new OrderBuilder().overLimit().build();
assertEquals(OrderResult.DECLINED, service.place(order));
The delta from raw construction is focus. The builder hides everything the test does not care about. So the only visible setup is the arrangement that causes the behavior under test. Records make this easier upstream with a compact constructor and validation. Similarly, the Mockito article’s small honest fakes play the same role for collaborators, such as an in-memory repository shared by tests. One caution applies, though. Shared fixture factories, the Object Mother pattern, become change-amplifiers when every test reaches through them. So keep builders small and per-test-class where possible.
The Flakiness Tax: Sleeps, Clocks, and Shared State
Flakiness, a test that passes and fails on the same code, is the single worst property a suite can have. Every flaky test trains the team to distrust the whole suite. The standard causes have standard cures, and the first one is the most common test smell in Java:
// WRONG: sleep-based waiting: too slow when it works, flaky when it is too short
Thread.sleep(2_000); // "give it time"
assertEquals(1, counter.get()); // passes on fast machines, flakes on CI
// RIGHT: determinism by injection
// - inject a Clock: tests pass the fixed instant, production passes Clock.systemUTC()
// - inject an executor: tests pass a same-thread executor, no scheduling race
// - use awaitility or a countDownLatch for genuinely async code, with a deadline
var fixed = Clock.fixed(Instant.parse("2026-01-01T10:00:00Z"), ZoneOffset.UTC);
var service = new LateFee(fixed); // the date is a constant, not a guess
The delta is who controls time and scheduling. The sleep delegates timing to the machine’s mood, while the injected Clock makes time an input like any other. The same principle covers the rest of the flaky family. Seed your randomness so failures reproduce, using the utility classes article’s Random with a fixed seed. Also, never depend on method execution order, which JUnit deliberately does not promise. Finally, test concurrent code with the downloader’s discipline, bounded waits with deadlines instead of sleeps. A test that can fail for no reason is a production bug pointed at your own pipeline.
The Anti-Pattern Catalog
| Anti-pattern | The symptom | The fix |
|---|---|---|
| Sleep-based waits | Slow suite, flaky on CI | Inject Clock and executors, use bounded waits with deadlines |
| Shared mutable state | Passes in one order, fails in another | @BeforeEach fresh state, no static fields in tests |
| Test interdependence | Deleting one test breaks others | Each test builds its own given, runs alone and in any order |
| Unseeded randomness | Failure cannot be reproduced | Fixed seed, logged in the failure output |
| Real I/O in unit tests | Minutes-long suite, network flakiness | Push down the pyramid, mocks and fakes for unit tier |
| Assertion-free tests | Green while nothing is verified | Every test asserts, or it is not a test |
| Mystery guests | Setup far from the assertion, magic values | Builders and local arrangement, name every constant |
| Testing implementation details | Refactors fail tests without behavior change | Assert observable behavior and contract interactions |
How Real Systems Do This
Healthy engineering organizations treat the test suite as a product whose users are the developers. They measure it like one, with suite runtime in the CI dashboard and flake rate tracked and quarantined per test. They also enforce a rule: fix or delete a quarantined test within a sprint. After all, quarantine without a deadline is deletion with extra steps. The pyramid’s ratios hold in practice at every mature shop I have worked with. Fast unit tests run on every save, the integration band on every push, and end-to-end on merge and nightly.
The flaky test that taught me the deadline rule is a story about ignoring the alarm until it mattered. A checkout integration test failed roughly one run in ten. Everyone assumed environment weather, and the routine became rerun until green.
Months later, a customer-visible duplicate charge came in. The investigation traced it to a real race in the code the flaky test had been imperfectly observing. In other words, the test was right one time in ten, and the nine green runs were the machine being lucky, not the code being correct. The post-incident rule at that company became the one I keep now. We investigate any test that fails intermittently the same week. A flaky test is a signal you have already paid for, and ignoring it means paying again in production. Most flakes turn out to be test bugs, the sleep and shared-state family. But the ones that are not test bugs are the most valuable failures a suite ever produces.
Decision Framework
- What tier does the verification need? The cheapest one that can honestly perform it: pure logic in unit tests, seam behavior in integration, critical journeys end-to-end.
- Is the setup wide or the data complex? A test builder with valid defaults, and a small fake for collaborators, before raw construction or a giant fixture factory.
- Does the code read the clock, schedule, or randomize? Inject those three, Clock, executor, seed, so the test controls them instead of racing them.
- Is the test asserting behavior or implementation? Behavior survives refactors, while implementation coupling fires on every harmless change. The fix is rewriting the assertion.
- Did a test fail intermittently? Investigate within the week: it is a test bug or a real bug, and both deserve the sprint.
- Is the suite slower than a coffee refill? Measure it, tag the tiers, push slow tests down the pyramid, and keep the fast tier under ten seconds.
- Is a test hard to name in one behavior sentence? Split it: the naming difficulty is the design telling you the test does two things.
When NOT to Use This
- Do not write integration tests for pure logic. The database adds minutes and flake to a verification JUnit can do in milliseconds, and the pyramid is the reason.
- Do not build fixture machinery for a three-field constructor. A builder for trivial data is ceremony, and direct construction reads better.
- Do not test the framework’s code, such as JUnit assertions, Mockito matchers, or collection internals. Instead, trust the toolchain and spend the effort on your logic.
- Do not chase 100 percent coverage as a goal. Coverage is a map of untested code, not a quality score. Also, the last few percent are often tests that verify nothing.
- Do not add tags, profiles, and layers for a suite that runs in five seconds. Process complexity must be earned by suite size.
Common Mistakes
- Letting flaky tests live: each one converts a red build from a signal into noise. Then the team’s response, rerunning until green, is the suite dying in slow motion.
- One giant test class per feature: thousand-line classes hide failures, and @Nested facets make the report navigable.
- Testing only the happy path: the exceptions article’s failure paths are the code most likely to be wrong. Yet they are the least likely to run by accident.
- Building the Object Mother into a change amplifier: every test touches one shared factory. So any fixture change becomes a suite-wide migration.
- Unmeasured suite runtime: nobody notices a suite creeping from 30 seconds to 12 minutes. Yet the fix gets cheaper the earlier the trend is visible.
- Mixing all tiers in one untagged run: developers stop running the full suite locally because the database tests need the database. So tag tiers and let the fast tier run everywhere.
- Deleting tests instead of fixing them on refactor: each deletion without a replacement removes a verified belief. Then the suite’s coverage story becomes a guess.
Key Takeaways
- The pyramid sets the ratios: many fast unit tests, a focused integration band, and a thin end-to-end layer. Each tier up costs speed and trust.
- Naming is documentation: behavior-stating method names, @Nested facets, and given, when, then make failures diagnosable in seconds.
- Fixtures and builders tame hard setup: valid defaults, fluent overrides, and the interesting value as the only visible arrangement.
- Flakiness is the suite killer: no sleeps, injected Clock and executors, seeded randomness, and order independence by construction.
- An intermittently failing test is a paid-for signal. Investigate within the week, because most are test bugs and the rest are real ones.
- Writing good tests means treating them as code with readers and a trust requirement. The suite is a product, measured, maintained, and believed.
- Push every verification to the cheapest honest tier. After all, the arithmetic of the pyramid is what keeps the suite fast enough to matter.
FAQ
What makes a good test in Java?
Writing good tests comes down to determinism, independence, and a name that states one behavior. A good test builds its own state and runs against a mock or fake at the boundary. It also asserts observable behavior and passes or fails the same way on every run. Speed is the fourth property, because a suite nobody runs protects nothing.
What is the test pyramid?
The ratio pattern for suites. It puts many fast unit tests at the base, fewer integration tests at the seams, and a thin end-to-end band at the top. Each tier up is slower and more failure-prone. So pushing verifications to the cheapest honest tier keeps the whole suite fast and trustworthy.
What is a test fixture in Java?
The known state a test runs against. It includes the object under test, its collaborators, and the data arranged in @BeforeEach or a builder. A good fixture is small, local to the test class, and shows only the values the behavior under test depends on.
How do I stop flaky tests?
Remove the three standard causes. Replace sleeps with an injected Clock and executors, and replace shared mutable state with fresh per-test setup. Also, replace unseeded randomness with fixed seeds. Then enforce the process rule: investigate every intermittent failure within the week, because it is a test bug or a real one.
What is a test builder pattern?
A small fluent class that constructs test data with valid defaults and one named override per interesting value. As a result, setup reads as one line that shows only the cause of the behavior under test. It replaces wide raw constructors and discourages shared fixture factories that amplify change.
Conclusion
Writing good tests has turned the suite into a product in your hands. It is pyramided for speed, named for diagnosis, fixed-up with builders and fakes, and protected from the flakiness that kills trust. From the JUnit unit to the Mockito boundary to this article’s suite engineering, the testing story of the course is complete. Every project from here ships with a suite that gets run because it deserves to be.
The next article covers the operational sibling of tests: logging with SLF4J and Logback. Logging is the structured, leveled record of what production did. It pairs with the suite as the second half of your observability story.
A suite that runs in seconds and never lies is infrastructure. Build that, and refactoring becomes safe enough to be fun.
Last updated on 13 September 2026.
