Java

Serialization in Java: How It Works and Why to Avoid It

Executive Summary

Serialization in Java starts with Serializable, a marker interface. Implement it and the runtime can flatten an object graph, via reflection over fields, into a byte stream. Later, it can reconstruct the graph, restoring fields directly without calling your constructors. Meanwhile, serialVersionUID is the compatibility contract between a class and its bytes. Pin it explicitly, or the compiler-generated value changes with any structural edit and old data throws InvalidClassException. Also, transient excludes fields from the bytes, the tool for derived values and secrets.

However, the case against has three arguments. The first is security, because deserialization runs before your validation and historically enabled remote code execution through gadget chains. The second is compatibility, because the bytes couple to private field layouts that any refactor can break. The third is API design, because Serializable spreads transitively to every field type. As a result, new code uses explicit formats instead: JSON across service boundaries, validated text formats inside tools, and JDBC for persistence. Where legacy deserialization cannot be removed, the required mitigation is JEP 290’s ObjectInputFilter.

How Serializable Works

Strangely, the interface has no methods. That is the first clue that the machinery lives elsewhere. The runtime uses reflection to walk your object graph field by field:

import java.io.*;

class User implements Serializable {
    private static final long serialVersionUID = 1L;  // the compatibility pin: declare it

    private final String name;
    private transient String cachedLabel;             // excluded from the bytes

    User(String name) { this.name = name; }
}

// object graph -> bytes
try (var out = new ObjectOutputStream(new FileOutputStream("user.ser"))) {
    out.writeObject(new User("Imraan"));
}

// bytes -> object graph
try (var in  = new ObjectInputStream(new FileInputStream("user.ser"))) {
    var back = (User) in.readObject();
}

Three mechanical facts matter more than the API details. First, writeObject serializes the whole reachable graph: the User, its fields, its fields’ fields, with cycle handling. As a result, one accidental reference to a large structure silently inflates the payload. Second, readObject reconstructs fields directly. Your constructors never run, so invariants you enforce in construction are not enforced on the way back in unless you write readObject validation yourself. Third, the stream embeds class metadata, name and serialVersionUID, and the reader refuses mismatches:

// the failure this pin prevents:
// deploy 1: class User { String name; }                // UID auto-generated: 482913...
// deploy 2: class User { String name; String email; }  // UID auto-generated: 715530...
// a blob written by deploy 1, read by deploy 2:

java.io.InvalidClassException: User; local class incompatible:
stream classdesc serialVersionUID = 482913..., local class serialVersionUID = 715530...

Without an explicit serialVersionUID, the runtime derives one from the class’s shape. So adding a field, reordering members, or sometimes just recompiling with a different toolchain changes it. Therefore, pin the value as a constant and the stream stays readable across compatible edits. The Serializable contract documents exactly which changes the platform considers compatible. Meanwhile, transient keeps derived values, caches, and secrets out of the bytes. On reconstruction it leaves those fields null or zero, so transient fields need lazy re-initialization, not assumptions.

Why the Industry Walked Away

The mechanism still works, ships in every JDK, and is documented in the serialization specification. However, the avoidance verdict rests on three arguments, each with a decade of production evidence:

Argument The mechanism The production consequence
Security readObject reconstructs objects before any validation of yours runs; classes with exploitable readObject implementations become gadget chains Historically, untrusted deserialization led to remote code execution; whole exploit toolkits target Java’s format specifically
Compatibility Bytes encode private field names, types, and layout Renaming a private field, or any structural edit, breaks stored data with InvalidClassException
API design Serializable is viral: every field’s type must be serializable too, and the serialized shape becomes public API Refactoring freedom dies; the wire format locks internals in place

Notice what the third row means for the domain types you design. Once User implements Serializable, so must Address, and every type Address holds. Then the UID and field layout of each become part of a contract you did not mean to publish. The security row’s severity is worth one more sentence. The JDK’s own response, JEP 290’s serialization filtering, exists because the format cannot be fixed, only fenced. The next section shows the fence.

The Alternatives: Explicit Boundaries

In practice, the recommended replacement is not one technology but one principle: serialize data you promised to serialize, in a format you designed, at a boundary you own. In other words, the purpose picks the format:

Purpose Use Why it wins
Data across services (wire format) JSON, Part 9’s HTTP and JSON stack Human-readable, cross-language, versionable by adding optional fields
Internal tool file format Explicit text format, the CatalogStore pattern Validated on construction, diffable in git, debuggable by eye
Relational persistence JDBC, Part 8’s data layer SQL types, queries, migrations; the database owns the format
Immutable data transfer Records mapped to JSON Compact carriers without the Serializable viral contract
Cross-JVM messaging (when you control both ends) A versioned binary format by contract Explicit schema, evolution rules, no class-shape coupling

For example, compare the wrong and right shape of the same decision:

// WRONG: implementing Serializable "just in case"
class Money implements Serializable {          // now every field type must be too,
    private BigDecimal amount;                // the UID joins your wire contract,
    private String currency;                  // and private fields become public API
}

// RIGHT: an explicit boundary format, the CatalogStore pattern
String line = "%s|%s|%s".formatted(id, amount, currency);
// text: typed by the writer, validated by the reader, versionable by design
// Money itself stays free of any serialization contract

The delta: the wrong version couples Money’s private shape to a byte format forever. By contrast, the right version puts a format at the edge, the CatalogStore already works this way, keeps validation in the reader’s constructor, and leaves the domain type free. This is the same lesson as the clone article: convenience mechanisms that bypass your constructors take invariants with them.

If You Inherit Serializable: Survival Guide

Legacy systems will hand you serialized sessions, caches, or blobs. Fortunately, the survival guide is short. First, pin serialVersionUID on every legacy class that has stored data. Next, mark derived and sensitive fields transient. Validate invariants after reconstruction, in a readObject method or after, because constructors do not run. And where the input is not fully trusted, install JEP 290 filters, the fence the JDK built for exactly this hole:

var in = new ObjectInputStream(untrustedBytes);

in.setObjectInputFilter(info ->
        info.serialClass() == Session.class
            ? ObjectInputFilter.Status.ALLOWED
            : ObjectInputFilter.Status.REJECTED);   // one class, nothing else

var session = (Session) in.readObject();

The filter runs before deserialization, checking class, depth, array sizes, and references against a policy you state. As a result, the gadget chains die at the door rather than inside your heap. If you can name the single class you expect, reject everything else, because that strictness is the point.

How Real Systems Do This

The modern stack converged years ago. HTTP APIs speak JSON, caches and queues speak their own versioned formats, and databases speak SQL. Meanwhile, Serializable appears mainly in legacy middleware and older frameworks. For example, Spring’s session and caching abstractions moved to explicit serialization strategies, and the big data tools use their own schemas. Also, the JDK team’s long-term direction has been to replace the mechanism’s unsafe parts rather than extend its use. As a result, when reading industry code, your expectation should be inverted from 2005: Serializable is a smell that earns an explanation, not a default that earns a pass.

The InvalidClassException story is one I have watched happen in production more than once. A team cached serialized user objects in a session store. Then a routine release added one field to the User class. Every existing session died on next contact with a local class incompatible exception, and users were logged out en masse mid-afternoon. No test caught it, because tests serialize and deserialize within the same class version. That is the exact scenario the pin and the explicit format exist to prevent. The fix that day was quick: pin the UID. However, the durable fix was replacing the blob cache with an explicit key-value representation the team could evolve deliberately. When you inherit serialized state, assume the same failure is scheduled for your next release, and pin or replace accordingly.

Decision Framework

  1. Are you designing a new type or a new wire format? Do not implement Serializable. Instead, choose JSON for service boundaries or an explicit format for tools, and keep the domain type contract-free.
  2. Does existing code hand you serialized bytes? Inventory every class with stored data and pin serialVersionUID on all of them before your next release, not after.
  3. Does any deserialization path touch untrusted input, user uploads, queues, caches writable by others? Install JEP 290 filters on those streams, allow-listing the expected classes, or remove the path.
  4. Do you control both ends and need Java-to-Java transport? Prefer a versioned explicit format, because Java serialization buys you nothing over it and costs the UID coupling.
  5. Are derived or sensitive fields in a class that must stay serializable? Mark them transient and recompute or re-request them after readObject.
  6. Are constructors enforcing invariants? Add readObject validation that re-enforces them, because reconstruction bypasses construction.

When NOT to Use This

  • Never deserialize input you do not fully control, such as user uploads, request bodies, and queue payloads, without a filter that rejects unexpected classes. After all, the historical result is remote code execution, not an exception.
  • Never add Serializable to a new domain type “for flexibility”. The viral contract spreads to every field type and freezes private structure into public API.
  • Never let the compiler generate serialVersionUID for data that outlives a process. Auto-generated UIDs change with class shape, so every change is a silent breaking release.
  • Do not cache serialized objects as a persistence strategy. The bytes are opaque, unqueryable, and unindexable, which is everything a database or key-value store exists to solve.

Common Mistakes

  • Forgetting serialVersionUID on a class that already has bytes stored somewhere. Then the next structural edit breaks every blob at once with InvalidClassException.
  • Serializing the wrong graph: for example, one reference to a parent object, a cache, or a whole service pulls megabytes into a stream meant for one small value.
  • Caching computed fields: values that are derived should be transient, or the stream resurrects stale computations alongside the data.
  • Writing secrets into bytes: passwords, tokens, and keys serialize like any other field. As a result, .ser files, caches, and payloads end up in places nobody audited.
  • Trusting the stream: treating readObject’s output as valid because it deserialized, when reconstruction ran no constructor and enforced no invariant.
  • Replacing serialization with another implicit format: hand-rolled byte hacks with DataOutputStream have the same class-shape coupling, minus the documentation.

Key Takeaways

  • Serializable is a marker interface, and the runtime flattens the reachable object graph via reflection and reconstructs it without running your constructors.
  • Pin serialVersionUID explicitly on any class whose bytes outlive the process, or every structural edit silently breaks stored data with InvalidClassException.
  • transient excludes derived and sensitive fields; those come back null or zero and need re-initialization.
  • The case against: deserialization ran before validation and enabled historical code execution. Also, bytes couple to private class shape, and the contract spreads virally through field types.
  • The replacement is a principle: explicit formats at boundaries you own. Examples are JSON across services, validated text formats in tools, JDBC for persistence, records as carriers.
  • Where legacy deserialization stays, JEP 290 filters are mandatory on untrusted input: allow the classes you expect, reject everything else.
  • Reconstruction bypasses construction: re-validate invariants after readObject, in code, every time.

FAQ

What is serialization in Java?

The runtime mechanism that converts an object graph into bytes via ObjectOutputStream.writeObject, and reconstructs it later via ObjectInputStream.readObject, marked by the Serializable interface and implemented through reflection over fields.

Why is Java serialization insecure?

readObject reconstructs objects before your validation runs, so crafted byte streams can trigger arbitrary code inside exploitable classes, historically causing remote code execution. The JDK’s mitigation is JEP 290 filtering, which checks classes before deserialization, and the industry’s mitigation is explicit formats instead.

What is serialVersionUID?

The compatibility pin between a class and its serialized bytes. Declared explicitly, it keeps old data readable across compatible class changes; left to the compiler, it is derived from class structure and changes with any edit, breaking stored blobs with InvalidClassException.

What is transient in Java?

A field modifier that excludes a field from serialization. Derived values, caches, and secrets are marked transient; on deserialization those fields come back null or zero and must be recomputed or re-requested.

What should I use instead of Java serialization?

An explicit format at a boundary you own: JSON across services, a validated text format inside tools, JDBC for persistence, or a versioned binary format when you control both ends. The domain type stays free of any serialization contract.

Conclusion

You now know both halves of serialization in Java. The mechanism is reflection over fields, bytes embedded with class metadata, and reconstruction without constructors. The verdict is explicit formats at owned boundaries instead. Serializable is not deprecated and will not disappear. However, every new design decision you make should treat it as legacy: readable, maintainable, fenced with filters, and never extended.

Next, the following article moves from bytes on disk to bytes on the wire: networking in Java, sockets, and the modern HTTP client. These are the fundamentals that Part 5’s concurrency work and Part 9’s REST stack both build on.

Serialize what you promised, in a format you designed, through constructors that validate. The mechanism that promised to do it for you was the lesson.

Last updated on 24 September 2026.

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *