Java

Garbage Collection in Java: Generations, Collectors, and Tuning

Executive Summary

Garbage collection in Java starts from one rule. An object is garbage when no path from a GC root reaches it, whether from thread stacks, static fields, or JNI references. The collector traces those paths rather than counting references. The heap is organized by age because of the weak generational hypothesis: most objects die young. New objects allocate in eden. Then, surviving a minor collection promotes them through survivor spaces into the old generation, where collection is rarer and costlier.

A GC pause is the stop-the-world moment when threads freeze at safepoints so the collector can act. As a result, collector choice is a pause-profile choice. Serial and Parallel suit batch and small heaps. Meanwhile, G1, the Java 9+ default per JEP 248, gives bounded pauses on most services. Finally, ZGC per JEP 333 delivers sub-millisecond concurrent pauses on large latency-critical heaps. Everything is observable with -Xlog:gc, whose pause lines show before, after, and duration. Also, the GC tuning guide documents every flag.

The tuning hierarchy that works has four steps. First, reduce the allocation rate, starting with unbounded caches and retained collections. Then size -Xmx to the workload and container. Next, pick the collector by pause budget, and only then change flags with evidence. Never call System.gc, never use finalizers, and never tune a number you have not seen in a log.

Reachability: How the Collector Decides You Are Done

The collector never counts references, despite the folklore. Instead, it traces. Starting from the GC roots, every local variable on every thread’s stack, every static field, and every JNI reference, it walks the object graph. Everything the walk reaches is alive. Everything the walk misses is garbage, regardless of how many references it had, because none of them were themselves reachable:

var book = new Book("Effective Java");   // reachable: 'book' is a local on this stack
book = null;                            // now: is anything else holding a path to it?
                                        // if not, the object is unreachable: garbage

// roots: thread stacks (locals), static fields, JNI global refs
//   everything reachable from a root: kept
//   everything else: reclaimed, no matter how it was used before

Two consequences of tracing are worth carrying into production work. First, a single forgotten reference retains an entire graph. The usual suspects are collections, caches, and listeners that outlive their interest. Second, the trace explains why the collector can move objects and why references must not be read mid-collection. That is what safepoints are for, as the pauses section below shows.

Two legacy mechanisms get their formal retirement here, because they are GC-adjacent and still appear in old code. System.gc() is a request the JVM may ignore, and calling it forces pauses you did not budget for. In fact, the flag to disable such calls exists, -XX:+DisableExplicitGC, and it appears in real deployments for exactly this reason. And finalizers, the Object class’s finalize method, are deprecated and dangerous: finalization can delay reclamation arbitrarily and resurrect objects. The modern alternative is try-with-resources from the exceptions article, or Cleaner for the rare genuine case, and nothing else.

Generations: Why the Heap Is Organized by Age

Decades of measurement support one observation, the weak generational hypothesis: most objects die young. Request objects, intermediate strings, iterators, the debris of a computation, all unreachable almost immediately. HotSpot exploits that with an age-structured heap:

  new objects
       v
   [ EDEN ]  --minor GC: cheap, frequent-->   [ SURVIVOR spaces ]
       |                                          |  survive a few cycles,
       |  most die here: collected fast             |  aging with each escape
       v                                          v
   reclaimed at near-allocation speed      [ OLD GENERATION ]
                                               rare, heavier collections
                                               (or: fully concurrent, in ZGC)

The mechanics are simple. Allocation happens in eden, and thread-local buffers make new cost a pointer bump. When eden fills, a minor collection marks and copies the survivors into a survivor space. As a result, it compacts the live set for free and leaves eden empty and reusable.

Objects that survive several minors are promoted to the old generation, which is collected far less often. After all, old objects are empirically the ones that stay. Consider the performance meaning of “minor” and “major”. Minor collections are short and frequent, the price of allocation. In contrast, old-generation collections, or full collections, are the expensive ones your p99 graphs notice. This whole structure is why allocation-heavy code shows up as GC pressure. Examples include boxing in hot loops, string concatenation in iterations, and churning intermediate collections. The cause is not that allocation is slow. Rather, every allocation is a future minor collection’s workload.

Garbage Collection in Java: The Collector Menu

The collector is the strategy that walks this structure. In particular, the four mainstream choices trade throughput against pause length in different ways:

Collector Strategy Pause profile Chosen for
Serial Single-threaded, stop-the-world Long on big heaps, fine on small ones CLI tools, tiny containers, single-core
Parallel Multi-threaded, stop-the-world Shorter than Serial, still full pauses Batch jobs where throughput beats latency
G1, default since Java 9 Region-based heap, concurrent marking, incremental evacuation Pauses bounded by -XX:MaxGCPauseMillis, default 200ms Most services: the safe default
ZGC Concurrent almost everywhere, colored pointers Sub-millisecond typical, independent of heap size Latency-critical services with large heaps
java -XX:+UseG1GC  -Xmx2g  -Xlog:gc  App      // the default, stated explicitly
java -XX:+UseZGC  -Xmx16g -Xlog:gc  App      // when p99 must stay flat

// -Xlog:gc turns on the unified GC log, JDK 9+:
//   [gc] GC(12) Pause Young (Normal) 45M->12M(256M) 3.812ms
//   read: before -> after (capacity), and the pause duration

The table hides some guidance. First, the JVM’s ergonomics already pick sensibly from machine class and heap size. Second, G1 is the default for the heap sizes services use. Therefore, collector changes are the third lever, not the first. The profiling article gives the decision evidence, and the tuning hierarchy below gives the order.

Pauses: What Actually Stops, and When

A GC pause is a stop-the-world event. Every Java thread freezes at a safepoint, a place the VM has marked safe for inspection. Meanwhile, the collector performs whatever phase needs a stable heap. Which phases stop the world depends on the collector. Serial and Parallel stop for everything. G1 stops briefly for incremental evacuation but marks concurrently. Finally, ZGC compacts concurrently too, pausing only for tiny root scans. What the pause means for your service fits in one sentence: every request in flight waits the pause length. That is why GC shows up as latency spikes at the tail, p99 and p99.9, while the median stays smooth.

The safepoint detail matters for diagnosis, because not every pause in a Java program is GC. For example, threads also stop at safepoints for other VM operations, and JIT deoptimization can stall. One discipline separates a real diagnosis from folklore. The -Xlog:gc flag shows you every GC pause with timestamps and durations. As a result, one look confirms or refutes a suspected GC problem. Also, the architecture article’s jstat -gc watches the heap rates live.

The Tuning That Actually Matters

Everything above reduces to a working order. It is worth stating as a hierarchy, because the industry’s instinct is to start at the bottom:

  1. Reduce the allocation rate. Unbounded caches, ever-growing collections, boxing in hot loops, churning strings: every avoided garbage object removes a future collection’s workload. This is the only step that makes every later step easier.
  2. Size the heap to the workload and the container. Keep -Xmx below the container limit, large enough that promotion fits and small enough that pauses stay bounded. Too-small heaps thrash, while too-big heaps pause longer per collection.
  3. Choose the collector by pause budget. G1 for most services, ZGC when p99 must stay flat on a large heap, Parallel for pure batch throughput.
  4. Then flags, with evidence. MaxGCPauseMillis, generation ratios, and their relatives, changed one at a time, measured against logs, per the profiling article.

Step one’s most common shape in services is the unbounded cache. It is worth seeing the WRONG and RIGHT side by side. After all, the wrong version looks like good engineering until load arrives:

// WRONG: a cache that only grows - every entry is reachable forever,
// promoted to old gen on arrival, and the collector can never reclaim any of it
private final Map<String, Price> cache = new ConcurrentHashMap<>();

// RIGHT: bounded and evicting - the cache holds the working set, not history
// (a size cap with LRU eviction: a small library, or a LinkedHashMap with
//  removeEldestEntry under a lock, or your cache library of choice, bounded)
private static final int MAX_ENTRIES = 10_000;   // a number, reviewed, configurable

The delta is reachability economics. An unbounded cache converts your heap into an archive. Then old-generation collections pay for the archive every cycle, whether or not the concurrent collection is otherwise perfect. The same logic indicts listener lists that never deregister and static maps that accumulate per request. It also indicts any long-lived reference to what should have been a short-lived object. In short, the collector cannot free what you still point to. Once the code is clean, turn to the language-neutral playbook in Garbage Collection Tuning for Backend Engineers. It covers heap headroom, container limits, and canary rollouts for collector flags.

How Real Systems Do This

Production defaults for garbage collection in Java have quietly converged. G1 with a right-sized heap serves the vast majority of services. ZGC appears where p99 flatness is a product requirement, trading some throughput for latency consistency. Meanwhile, Parallel survives in batch and data-processing jobs where a long pause is invisible. Deployment pipelines turn on -Xlog:gc everywhere and ship the logs to monitoring. As a result, the GC conversation is never “it feels slow”. Instead, it is “old-gen collections went from 2 to 40 per hour after Tuesday’s deploy”. The flags teams actually maintain are few. They cover heap size, the collector, a pause goal, and a metaspace bound for class-generating workloads.

My favorite GC war story is a humbling one, because the bug was mine. A pricing service showed textbook symptoms: p50 fine, p99 ugly in periodic spikes, every dashboards’ first suspect GC. The first fix attempt was tuning, MaxGCPauseMillis and heap ratios, and it changed nothing. The problem was not the collector’s speed but the garbage it was handed. A pricing batch I had written built a fresh multi-megabyte map of the entire catalog per request. Then the whole structure died at request end. In effect, the service allocated its own heap every few minutes.

The log made it obvious: minor collections of hundreds of megabytes at exactly the spike cadence. Therefore, the fix was algorithmic: compute once, share the immutable result, stop manufacturing garbage. Latency spikes gone, and the flags I had tuned were reverted untouched. That incident taught me the hierarchy this article leads with, in the order I violated it. Allocation comes first, sizing second, collector third, and flags last. Also, every step begins by reading the log instead of assuming the diagnosis.

Decision Framework

  1. Is there a latency problem at the tail? Confirm GC first with -Xlog:gc: pause times and cadence, before touching anything else.
  2. Are collections large and frequent? Reduce allocation: unbounded caches, boxing in hot loops, churning intermediates, per-request maps that should be shared immutable results.
  3. Is the heap sized to the workload and container? -Xmx under the container limit, with headroom for metaspace and stacks, and a promotion rate that fits between collections.
  4. Is there a pause budget the default G1 does not meet? Try MaxGCPauseMillis first, then ZGC when flat p99 justifies its throughput cost.
  5. Is the workload batch, latency-irrelevant? Parallel collector, throughput-first, and accept full pauses.
  6. Is something retaining memory? A heap histogram, jcmd GC.class_histogram, names the owners. Retention is a lifecycle bug to fix, not a heap to enlarge.
  7. Are you about to change a flag? Change one at a time, with before-and-after logs, and revert what did nothing. That way, the flag file stays a record of evidence.
  8. Did someone call System.gc or leave a finalizer? Remove both: explicit GC requests are pauses on demand, and finalization is deprecated machinery the exceptions article replaced.

When NOT to Use This

  • Do not tune any GC flag without a log line that motivated it. After all, unmeasured tuning is how flag folklore spreads, and the profiling article is the standard for evidence.
  • Do not enlarge the heap to fix a leak. It converts a visible failure into a slower one, and the histogram still names the owner in the end.
  • Do not switch collectors to treat a symptom the allocation rate caused. In other words, ZGC with a garbage fire is a faster collector for the same fire.
  • Do not reach for generation-ratio flags on G1 and ZGC. Modern collectors manage their own geometry. Indeed, the old-generation ratios belong to the era before region-based heaps.
  • Do not add System.gc calls for tidiness: you cannot manage memory politely this way, you can only force pauses.
  • Do not hold onto this article’s details as trivia to apply everywhere. In fact, most services run the default, sized right, and never need more than the log watch.

Common Mistakes

  • Assuming GC causes every latency spike. Confirm with -Xlog:gc first, because safepoint and JIT stalls imitate GC. The log separates them.
  • Unbounded caches and ever-growing collections: perfect thread safety plus no eviction equals a memory archive the collector cannot touch.
  • Calling System.gc in production code. It is a pause on demand, usually from a misunderstood “memory is low” instinct, and usually disabled by flag.
  • Finalizers for cleanup. They are deprecated, with delayed reclamation and resurrection risk. Instead, try-with-resources replaced them, and the course has used it since the file tools.
  • Confusing minor with full collections in a log. Minors are the price of allocation, while fulls are the ones your latency graph feels. So reacting to minors wastes effort.
  • Setting the heap above the container limit. The collector will lose the race with the container runtime, and the process dies with no Java diagnostics.
  • Blaming the collector for retention bugs. The collector is not slow. Instead, something is being kept alive, and the histogram says what.
  • Benchmarking GC changes without warm-up and steady load. After all, the first minutes of any JVM measure its cold start, not its collection behavior.

Key Takeaways

  • The collector traces reachability from GC roots: thread stacks, statics, and JNI refs. Unreachable objects are garbage no matter how many references they once had.
  • The heap is generational because most objects die young. It uses eden allocation, survivor aging, and old-generation promotion. Minors are cheap and frequent, while full collections are expensive and rare.
  • A GC pause is a stop-the-world safepoint event, and collector choice is pause-profile choice. Serial and Parallel suit batch and small heaps, and G1 is the service default. Finally, ZGC suits sub-millisecond pauses on large latency-critical heaps.
  • -Xlog:gc is the ground truth, logging every pause with before, after, capacity, and duration. Therefore, no GC diagnosis should skip it.
  • The tuning hierarchy: reduce allocation first, size the heap second, choose the collector third, flags last with evidence.
  • Unbounded caches are the classic service GC problem: reachability defeats collection, and bounds restore it.
  • System.gc is a pause on request and finalization is deprecated: try-with-resources and Cleaner are the cleanup story.
  • Most GC problems are allocation problems wearing a collector’s clothes. Fix what you hand the collector before you reconfigure the collector.

FAQ

How does garbage collection work in Java?

The collector periodically traces the object graph from GC roots: thread stacks, static fields, and JNI references. Then it reclaims everything unreachable. HotSpot’s generational heap collects the young generation cheaply and often, because most objects die young. In contrast, it collects the old generation rarely. Also, the collector’s phases run partly concurrent and partly at stop-the-world pauses.

What are GC roots in Java?

The starting points of the reachability trace. They include local variables on live thread stacks, static fields of loaded classes, and JNI global references. An object survives if a path from some root reaches it, and reference counts play no role in the decision.

What is the difference between G1 and ZGC in Java?

G1, the default since Java 9, divides the heap into regions and marks concurrently. Then it evacuates in incremental stop-the-world pauses bounded by a pause goal, typically tens of milliseconds. In contrast, ZGC, JEP 333, does nearly everything concurrently. It delivers sub-millisecond pauses largely independent of heap size, trading some throughput for latency consistency.

What is a GC pause in Java?

A stop-the-world interval where all Java threads freeze at safepoints while a collector phase runs. Pause length is what latency-sensitive services feel at the p99 tail. That is why teams choose collectors by pause profile. Also, -Xlog:gc records every pause with its duration.

How do I read GC logs in Java 21?

Enable with -Xlog:gc and read the pause lines. For example, take [gc] Pause Young (Normal) 45M->12M(256M) 3.812ms. It means a young collection moved the heap from 45 to 12 megabytes of a 256-megabyte capacity. It also stopped the world for 3.8 milliseconds. Cadence and duration of those lines, young versus old, are the diagnostic core.

Conclusion

You now hold the complete memory story of the JVM, including garbage collection in Java. Reachability is the death criterion, and generations exploit object lifetimes. Collectors are pause-profile strategies, and logs are the only honest source about all of it. The tuning hierarchy, allocation, sizing, collector, flags, is the operating procedure. Next, the profiling article supplies the measurement discipline that makes every step of it evidence-driven.

The next article completes Part 6 with profiling and performance tuning basics. It applies the scientific method to a running JVM, from reproducing the problem to measuring before and after. Finally, it covers jstat, jcmd, flight recorder, and profilers, the tools that turn complaints into numbers and numbers into fixes.

Hand the collector less garbage, and it will hand you better latency. Everything else is refinement.

Last updated on 9 September 2026.

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *