Java

Virtual Threads in Java: The New Concurrency Model

Executive Summary

Virtual threads in Java are threads whose stacks live on the heap and whose scheduling the JVM owns. Each one runs on a carrier platform thread. When it blocks on I/O, it unmounts and frees the carrier instead of parking an OS thread, per the virtual threads guide. The result: blocking code with asynchronous throughput, and a million concurrent waits costing heap, not OS threads. Create them with Thread.ofVirtual().start or, the production form, Executors.newVirtualThreadPerTaskExecutor. That means one virtual thread per task and no pool, so the link checker’s pool-sizing decision evaporates.

However, two things do not change. First, thread safety: the race conditions and locks articles apply identically, since virtual threads are still threads. Second, downstream limits matter, because a million concurrent tasks can still overwhelm a 20-connection database pool. That is why Semaphores bound shared resources where thread pools once did. The rules are short. Never pool virtual threads, because they are cheap and one-shot by design. Do not use them for CPU-bound work, which still needs cores. Finally, avoid pinning: a synchronized block around blocking I/O holds the carrier and restores platform-thread economics. Instead, use ReentrantLock where blocking happens under a lock. For I/O-bound services, virtual threads are the biggest simplicity-per-line win in modern Java.

The Old Trade-Off, and Its Price

To feel what changed, state the pre-Java-21 problem precisely. A platform thread is an OS resource with a fixed stack, roughly a megabyte, reserved whether used or not. As a result, a service could afford thousands, not millions. Every request that needed the network therefore cost one parked OS thread while it waited:

// the platform-thread world: waiting is expensive

  request ----> [platform thread, ~1MB] ----blocking I/O----> parked OS thread
                                                        10_000 concurrent waits
                                                        = 10 GB of stacks, OS limit hit

// the async workaround: waiting is cheap, code is not

  request ----> supplyAsync(...).thenApply(...).thenCompose(...)   // scales,
                                                                  // but the program
                                                                  // became pipelines,
                                                                  // callbacks, and
                                                                  // common pools

The CompletableFuture article is the second picture. It is a good tool that exists to paper over the first picture’s cost. The industry spent a decade writing reactive pipelines precisely because blocking was unaffordable at scale. As a result, the complexity you paid, in debugging, in onboarding, in operators that hide control flow, was the price of OS threads being heavy.

What a Virtual Thread Actually Is

The mechanism is a mounted/unmounted execution over a small pool of carrier platform threads. Each stack lives in heap objects that grow and shrink as needed:

   carrier platform threads (a few, ~ core count)
        |  mount: run a virtual thread's code on the carrier
        v
   [ v-thread 1 ][ v-thread 2 ][ v-thread 3 ]...[ v-thread 1_000_000 ]
        ^
        |  unmount: on blocking I/O, the v-thread detaches, its stack goes
        |  to the heap, the carrier picks up another v-thread immediately
        blocking wait costs HEAP (a small object), not an OS thread
Platform thread Virtual thread
Managed by The OS The JVM
Stack Fixed, ~1MB, reserved Heap-allocated, grows as needed
Practical count Thousands Millions
Blocking on I/O Parks an OS thread Unmounts, frees the carrier
Correctness rules The threads and locks articles Identical: same rules, same races
Pooling Mandatory, per the executor article Never: they are cheap and one-shot
Best workload CPU-bound, long-lived Blocking I/O, short-lived, high volume

The row worth reading twice is correctness. Virtual threads do not make shared state safe. Instead, they make it affordable to have many threads sharing it. Every race, every lock, every concurrent collection from this part of the course applies unchanged. That is why this article could come after those rather than instead of them.

Using Virtual Threads in Java: One Line of Change

The API is deliberately the Thread and ExecutorService you already know, per the Thread API:

// one-off creation, the rare form
var vt = Thread.ofVirtual().name("fetch-1").start(() -> fetch(uri));

// the production form: an executor that spawns a virtual thread per task
try (var executor = Executors.newVirtualThreadPerTaskExecutor()) {
    var tasks = urls.stream()
            .<Callable<Result>>map(uri -> () -> fetchWithRetry(uri))
            .toList();

    var results = new ArrayList<Result>();
    for (var future : executor.invokeAll(tasks)) {
        results.add(future.get());            // all done: collect, exactly as before
    }
    System.out.println(report(results));
}

Compare that with the executor article’s link checker and count the differences: the factory method and nothing else. There is no pool size, no workload analysis, and no cores-versus-downstream debate, because there is no pool to size. A new virtual thread per task is the whole policy. Meanwhile, the blocking fetches cost heap instead of OS threads. The sizing question moved from threads, which are now free, to the shared resources downstream, which are not. That shift is the subject of the next section.

Bounding Resources, Not Threads

Removing the thread limit does not remove the limits. Downstream resources still saturate: a 20-connection database pool, a rate-limited third-party API, a socket the remote side will close. With platform threads, the pool accidentally bounded concurrency toward those resources, often wrongly but at least boundedly. Virtual threads unbound it completely, so the bounding must become explicit, and the tool is the Semaphore:

var dbGate = new Semaphore(20);          // the JDBC pool is 20 connections: match it

try (var executor = Executors.newVirtualThreadPerTaskExecutor()) {
    for (var book : queue) {
        executor.submit(() -> {
            dbGate.acquire();            // bound the RESOURCE, not the threads
            try (var conn = dataSource.getConnection()) {
                update(conn, book);      // blocking JDBC, cheap on a virtual thread
            } finally {
                dbGate.release();
            }
        });
    }
}
// 10_000 tasks, 20 at the database, the rest waiting on the semaphore: the
// concurrency the pool used to impose by accident, now imposed by design

This is the conceptual shift in one pattern. Platform-thread systems sized the threads and hoped the resources followed. In contrast, virtual-thread systems free the threads and size the resources directly. The semaphore count is a decision you can state, review, and change. That makes it better than the pool-size decision it replaces, because it names the actual constraint.

Pinning: The synchronized Exception to Cheapness

Virtual threads unmount when they block, with one important exception on Java 21. While a thread is inside a synchronized block or method, it stays pinned to its carrier. As a result, blocking I/O inside the region parks the carrier exactly as in the platform-thread world. The fix is the locks article’s ReentrantLock, which virtual threads can unmount under:

// WRONG: synchronized around blocking I/O - pins the carrier OS thread
synchronized (lock) {
    var body = httpClient.send(request, BodyHandlers.ofString()).body();   // blocks:
}                                                                            // carrier parked,
                                                                             // platform economics back

// RIGHT: ReentrantLock - the virtual thread unmounts while waiting for the lock
// and while blocked on the I/O inside
private final ReentrantLock lock = new ReentrantLock();

lock.lock();
try {
    var body = httpClient.send(request, BodyHandlers.ofString()).body();
} finally {
    lock.unlock();
}

The rule is narrower than the WRONG code suggests. A synchronized block around fast, in-memory work is completely fine. However, on Java 21 the lock wait pins too. A virtual thread waiting to enter a contended synchronized monitor cannot unmount either. So waiting for the monitor and blocking inside the held region both tie up the carrier. ReentrantLock, by contrast, lets the thread unmount in both cases. JDK 24 removes this limitation through JEP 491, so synchronized no longer pins there. When you suspect pinning, Java 21 can log pinned threads with -Djdk.tracePinnedThreads=full. Meanwhile, the JDK’s own synchronized regions have been steadily migrated. The habit to build now is the one the locks article already taught for other reasons. Keep lock regions short, and no lock region should contain a network call.

When They Win, and When They Do Not

Virtual threads are a throughput tool for waiting, not a speed tool for computing. The wins and non-wins divide cleanly:

  • Wins: high-volume blocking I/O, HTTP services, JDBC calls, message consumption, and file processing. They win anywhere concurrency equals the number of concurrent waits, especially in codebases with existing blocking libraries that async frameworks could never adopt.
  • No difference: CPU-bound work, parsing, hashing, compression, which still needs cores. Therefore, platform pools sized at availableProcessors remain correct, per the executor article’s sizing rules.
  • Actively wrong: pooling or caching virtual threads, which reintroduces the scarcity they removed. The same goes for ThreadLocal-heavy legacy code, whose per-thread caches can balloon into thousands of copies when the thread count grows a hundredfold.

How Real Systems Do This

Adoption is moving faster than any Java feature in recent memory, because the change is small and the code is already written for it. For example, Spring Boot 3.2 exposes a single property, and Tomcat serves requests on virtual threads. Also, every blocking library, JDBC drivers, HTTP clients, file APIs, works on virtual threads unmodified. That was the design’s entire point: adopt blocking code you already have, not a new programming model. The pattern you will see in production services in the Java 21 era has four parts. Request handling runs on virtual threads, and a platform pool takes CPU-bound background work. Explicit semaphores guard shared downstream resources. Finally, ReentrantLock appears where old code held synchronized across I/O.

My migration story is the pinning incident, and it doubles as the review lesson. A service moved its request handling to virtual threads and throughput rose exactly as hoped, for two days.

Then the p99 latency began climbing, and the thread dumps told the story. Carrier threads sat pinned inside a legacy synchronized method that wrapped a vendor HTTPS call. That meant 300 milliseconds of blocking per invocation, all of it parking carriers. The service had virtual threads on the outside and platform economics on the inside. The fix was mechanical: ReentrantLock with the same try-finally idiom the locks article taught. Moreover, the fix’s review rule is the one I now hand to teams. When you enable virtual threads, grep for synchronized methods that make network or database calls. After all, every one of them is a carrier waiting to be parked. The feature did not create that bug. Instead, it exposed a bad habit the platform used to make affordable.

Decision Framework

  1. Is the workload dominated by blocking I/O, HTTP, JDBC, files, sockets? Virtual threads per task, the default answer in modern Java.
  2. Is it CPU-bound, with real computation per task? Platform pool at availableProcessors; virtual threads add nothing to compute.
  3. Does a shared downstream resource have a hard limit, connection pool, API quota? Bound it with a Semaphore sized to the limit, explicitly, instead of letting thread counts bound it accidentally.
  4. Are there synchronized regions containing blocking calls? Migrate them to ReentrantLock before enabling virtual threads, or run with tracePinnedThreads and fix what it reports.
  5. Is per-request concurrency high and bursty, like web traffic? Virtual threads excel: a burst creates threads, not queues, and the burst drains as threads finish.
  6. Is the tool short-lived and sequential, a CLI run over a dozen items? Plain blocking code on the main thread, no threads of any kind: virtual threads are for concurrency you actually need.
  7. Does the code lean hard on ThreadLocal caches? Audit before scaling thread counts a hundredfold, or the cache multiplies with them.
  8. Are you migrating an async pipeline? Only if the pipeline exists purely for I/O scalability. In that case, virtual threads let the straight-line blocking version replace it, and that rewrite is usually a deletion.

When NOT to Use This

  • Do not pool virtual threads, ever. They are cheap and disposable by design. In fact, a pool reintroduces the scarcity the feature removed while adding queue management nobody needed.
  • Do not reach for them on CPU-bound work. Parsing, hashing, and compression scale with cores. As a result, a million virtual threads on eight cores is a million threads waiting for a turn.
  • Do not treat them as a substitute for bounded downstream resources. Unbounded thread concurrency against a bounded database pool is a queue in disguise. Only now, the queue items are threads holding memory.
  • Do not enable them on code with synchronized-over-I/O until it is audited. Otherwise, the pinning will cap you exactly where platform threads did. It also adds the confusion of virtual threads that behave like platform ones.
  • Do not rewrite working, understood async pipelines overnight. CompletableFuture code that is correct and monitored can migrate incrementally. Also, the practice article shows the straight-line version side by side for the link checker.

Common Mistakes

  • Pooling virtual threads: the single most common conceptual error, born from platform-thread habits; pool nothing, create per task.
  • Expecting faster single-threaded execution. Virtual threads change how many can wait, not how fast one computes. So latency per operation is unchanged or marginally different.
  • synchronized around network or database calls: pins carriers, restores megabyte-per-wait economics, and hides behind a working system until load arrives.
  • Forgetting downstream bounds. A million tasks against twenty connections is a million waiting threads and a connection pool under siege. That is why the Semaphore is part of the migration, not an optimization.
  • ThreadLocal caches on virtual threads: per-thread state multiplies with thread count, and a cached 10MB object becomes gigabytes at scale.
  • Using them for low-volume tools. A link checker over twenty URLs gains nothing. After all, the decision belongs to the concurrency you actually have.
  • Assuming thread safety changed. Races, locks, and concurrent collections apply identically. In fact, a million threads makes a race a million times more likely to fire, not less.

Key Takeaways

  • Virtual threads in Java break the blocking-versus-async trade-off. They are JVM-scheduled threads with heap-based stacks that unmount on I/O, so blocking code scales to millions of concurrent waits.
  • The production form is Executors.newVirtualThreadPerTaskExecutor: one cheap thread per task, no pool, no sizing. As a result, the link checker’s concurrency became one changed line.
  • Correctness is unchanged: races, locks, and concurrent collections apply identically, because a virtual thread is still a thread.
  • Bound resources, not threads: Semaphores around connection pools and quotas replace pool sizing, and name the real constraint explicitly.
  • Never pool virtual threads: cheap and disposable is the design, and pooling reinstates the scarcity the feature removed.
  • Pinning is the one performance trap on Java 21. Synchronized regions containing blocking I/O park carriers. ReentrantLock is the fix, and tracePinnedThreads finds the offenders.
  • CPU-bound work gains nothing: it still scales with cores, and platform pools at availableProcessors remain correct there.
  • Adoption is a property flag plus an audit, not a rewrite: blocking libraries work unchanged, which was the point.

FAQ

What are virtual threads in Java?

They are JVM-managed threads, finalized in Java 21 through JEP 444. Their stacks live on the heap, and they mount on a few carrier platform threads, unmounting when they block on I/O. Blocking waits cost heap instead of OS threads, so millions of concurrent blocking tasks become practical.

How do I create a virtual thread in Java?

One-off with Thread.ofVirtual().start(runnable), or in bulk through Executors.newVirtualThreadPerTaskExecutor, which creates a fresh virtual thread per submitted task. The bulk executor is the production form, since it plugs into the ExecutorService idioms you already use.

Should I pool virtual threads?

No, never. They are cheap, one-shot, and designed to be created per task, and pooling reintroduces the scarcity the feature exists to remove. Bound the downstream resources they touch with semaphores instead of bounding the threads.

What is pinning in virtual threads?

On Java 21, a virtual thread inside a synchronized block, or waiting to enter one, cannot unmount. So blocking I/O in that region parks its carrier platform thread, restoring platform-thread economics. The fix is ReentrantLock in the lock-try-finally idiom, and synchronized over fast in-memory work remains fine. JDK 24 removes this pinning through JEP 491.

What is the difference between virtual and platform threads?

Platform threads are OS threads with fixed megabyte stacks, practical in the thousands, and must be pooled. Virtual threads are JVM-scheduled with small heap-based stacks, practical in the millions, never pooled, and best for high-volume blocking I/O, while CPU-bound work still belongs on platform pools sized to the core count.

Conclusion

You now hold the feature that retires the hardest trade-off in server Java. It delivers blocking code that scales, adopted through one executor factory and a resource-audit. The rules are few: free the threads, bound the resources, and keep locks off the blocking paths. Meanwhile, everything this part taught about safety applies without modification.

The next article completes the modern concurrency story with structured concurrency. That API gives concurrent subtasks the lifecycle of an ordinary method call. As a result, a parent’s cancellation, failure, and completion propagate to its children automatically. Also, the link checker’s gather step stops being manual bookkeeping.

Write it blocking, run it scaled. The platform spent twenty-five years owing you both.

Last updated on 2 September 2026.

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *