Java

Practice: Build a Concurrent Downloader in Java

Executive Summary

The concurrent downloader you build here is ConcurrentLinkCheck, in two forms over the unchanged Part 4 domain: sealed Result, fetchWithRetry, report. Form 1 is the executor: a fixed pool of 8 threads with invokeAll collecting Futures. It inherits the three-step shutdown from try-with-resources and reports progress through an AtomicInteger incremented at each task’s end. In contrast, Form 2 swaps one line, newVirtualThreadPerTaskExecutor, making every fetch its own cheap thread and the pool-size decision disappear.

Both forms also add the politeness layer the Part 4 version lacked: a per-host Semaphore allowing at most 2 concurrent requests per hostname. A ConcurrentHashMap stores the gates via computeIfAbsent, the concurrent collections article’s atomic memoization pattern doing real work. Wrap each fetch in the gate’s acquire and release, and time both versions with System.nanoTime. Then record what the wall clock says: sequential time versus pool time versus virtual-thread time on the same URL list. Finally, the checklist verifies every behavior including the failure path. The extensions point toward the structured-concurrency form and CI-ready exit codes.

The Brief

Requirements, stated as behaviors you can verify:

Requirement Detail Exercises
Fetch all URLs concurrently Every URL in urls.txt, same Result classification as Part 4 ExecutorService.invokeAll, Callable
Bound overall concurrency Form 1: pool of 8. Form 2: virtual threads, unbounded threads but resources gated Pool sizing, virtual threads
Politeness per host At most 2 concurrent requests to any single hostname Semaphore per host, computeIfAbsent
Live progress completed/total printed as tasks finish AtomicInteger, thread-safe counters
Clean shutdown try-with-resources; the tool exits when work is done, never hangs Executor close semantics
Honest timing Wall-clock time for sequential Part 4 vs both concurrent forms System.nanoTime, measurement

The Plan

The concurrent downloader’s architecture is Part 4’s pipeline with a concurrent middle. Domain types unchanged, engine swapped, one new coordination piece:

  urls.txt
     |
     |  Files.readAllLines + parse (unchanged from Part 4)
     v
  List<URI>  ---------------------->  Invalid (unchanged)
     |
     |  for each URI: a Callable<Result> task
     v
  +--------------------------------------------------+
  |  CONCURRENT ENGINE (this practice's change)       |
  |    Form 1: fixed pool(8) + invokeAll             |
  |    Form 2: virtual-thread-per-task executor      |
  |                                                  |
  |    per task:  hostGate.acquire()                 |
  |              fetchWithRetry(uri)    (unchanged)  |
  |              hostGate.release()                  |
  |              progress.increment()                |
  +--------------------------------------------------+
     |
     v
  List<Result>  -->  report() (unchanged)  -->  report.txt + console

Three decisions made now save rework later. First, the domain types do not change. Concurrency wraps the outside, per the pools article, and the sealed Result hierarchy keeps the report exhaustive. Second, politeness is a per-host gate, not a global slowdown. So different hosts proceed in parallel while the same host is paced. Third, the engine swap, Form 1 to Form 2, should touch one line. If it touches more, the boundary leaked and the design wants another look.

Form 1: The Executor Version

The full engine, with progress, built on everything from the pools article:

import java.net.URI;
import java.util.ArrayList;
import java.util.List;
import java.util.concurrent.*;
import java.util.concurrent.atomic.AtomicInteger;

class ConcurrentLinkCheck {

    private final AtomicInteger completed = new AtomicInteger();   // thread-safe progress

    List<Result> run(List<URI> urls) throws InterruptedException {
        var hostGate = new HostGate();                              // politeness, next section

        try (var pool = Executors.newFixedThreadPool(8)) {          // 8 threads, bounded

            var tasks = urls.stream()
                    .<Callable<Result>>map(uri -> () -> {
                        hostGate.acquire(uri);                       // at most 2 per host
                        try {
                            return fetchWithRetry(uri);             // Part 4, unchanged
                        } finally {
                            hostGate.release(uri);
                            System.out.printf("\r%d/%d done",
                                    completed.incrementAndGet(), urls.size());
                        }
                    })
                    .toList();

            var results = new ArrayList<Result>();
            for (var future : pool.invokeAll(tasks)) {              // submit all, wait all
                results.add(getQuietly(future));                    // unwrap, below
            }
            return results;
        }   // close: waits for tasks, releases the threads. No shutdown ceremony.
    }

    private Result getQuietly(Future<Result> future) {             // pools article:
        try {                                                       // exceptions live in Futures
            return future.get();
        } catch (ExecutionException e) {
            return new Failed(null, "task failed: " + e.getCause());
        } catch (InterruptedException e) {
            Thread.currentThread().interrupt();
            throw new IllegalStateException("interrupted while collecting", e);
        }
    }
}

Notice how much of this is the pools article verbatim. The fixed pool, invokeAll batch, Future collection, ExecutionException unwrapping, and interrupt etiquette all come from it. The only genuinely new structure is the gate, and the gate is the next section’s single idea. Also, the try-with-resources close performs the graceful shutdown for you. As a result, the tool cannot hang or leak threads by construction.

Politeness: The Per-Host Gate

The Part 4 tool fetched one URL at a time, so it could not abuse anyone. Concurrent, it can. After all, 50 requests to one host in the same instant is the denial-of-service shape the networking article warned about. The gate bounds per-host concurrency, and the concurrent collections article’s computeIfAbsent builds one gate per host on demand:

class HostGate {
    private final ConcurrentHashMap<String, Semaphore> gates = new ConcurrentHashMap<>();

    void acquire(URI uri) throws InterruptedException {
        gateFor(uri).acquire();                 // blocks until this host has capacity
    }
    void release(URI uri) {
        gateFor(uri).release();
    }
    private Semaphore gateFor(URI uri) {
        return gates.computeIfAbsent(                       // atomic: exactly one
                uri.getHost(),                              // Semaphore per hostname
                host -> new Semaphore(2));                   // at most 2 per host
    }
}

Read the design against the tools it combines. First, computeIfAbsent guarantees one semaphore per host even when fifty tasks race to create the same gate. Next, the Semaphore’s acquire blocks politely instead of failing. Finally, release sits in the finally block where the exceptions article put every cleanup obligation. Consider eight pool threads with two slots per host. Four different hosts run at full pace, while a fifth host’s requests queue at its own gate. As a result, no single destination ever sees a burst.

Form 2: The Virtual Thread Version

Now the part’s payoff, and you should time the edit: swap one line, change nothing else, run again.

// Form 1:                                     // Form 2:
try (var pool =                                 try (var pool =
        Executors.newFixedThreadPool(8)) {              Executors.newVirtualThreadPerTaskExecutor()) {
                                                // ^ the ONLY change. Every task now
    ...exactly the same code...                 //   gets its own cheap virtual thread,
}                                               //   and the pool-size decision is gone.

Everything else survives untouched because the boundary held. The gate still paces each host, and the progress counter still climbs. Also, invokeAll still collects the same Futures, and the report still pattern-switches the same sealed types. The sizing question changed shape entirely. Form 1 bounded the work with 8 threads, a blunt instrument that also serialized unrelated hosts. In contrast, Form 2 removes the thread bound and leaves the politeness bound as the only limit. That is the semantically correct one, since 2-per-host was the constraint you actually meant.

Which form should ship? The measurements decide, and the next section makes you take them. However, the priors are worth stating. For a tool like this, tens to hundreds of URLs, both forms perform identically at human timescales. So the tie-breaker is taste and platform. Virtual threads are the direction the language is going. Also, the pool size in Form 1 is one more magic number to defend, while Form 2’s code is Form 1’s code minus a decision. In services, at thousands of concurrent waits, the virtual-thread form is the one that scales without re-plumbing.

Measure the Concurrent Downloader: The Honest Numbers

Timing harness first, one line around each run, per the blocking client code Part 4 built:

long start = System.nanoTime();
var results = check.run(urls);
long millis = (System.nanoTime() - start) / 1_000_000;
System.out.printf("%nchecked %d urls in %d ms%n", results.size(), millis);

Run the same URL list three ways: the Part 4 sequential tool, Form 1, and Form 2. Then write down what you see. Illustrative shape of the numbers on a 50-URL list spread across 10 hosts:

Version Typical wall time Why
Sequential (Part 4) The sum of all fetches One thread, one wait at a time
Form 1: pool of 8 Roughly the sum divided by 8, floored by the slowest host’s gate 8 lanes, shared
Form 2: virtual threads Roughly the slowest single fetch, floored by the slowest host’s gate Every wait concurrent, politeness binds

Your exact numbers will differ, since network variance guarantees it. However, the differences are the lesson, not noise. Concurrency turns a sum into a maximum, and the gate you added turns the maximum into a principled trade. Also watch what happens with a URL list that is mostly one host. All three versions converge, because the per-host semaphore, not the threads, is the binding constraint. That observation is the virtual threads article’s conceptual shift, arrived at by stopwatch.

How Real Systems Do This

This concurrent downloader is not a toy with production paint. Instead, it is the skeleton of real operational tooling. For example, web crawlers are concurrent fetchers with per-host politeness, exactly this design at larger scale. Similarly, CI link checkers are this tool with exit-code gating. Meanwhile, data ingestion services are producer-consumer pipelines with bounded queues, the same backpressure thinking applied to a different shape. The two decisions this practice forces, where to bound and how to be polite, are the two decisions every production fetcher makes. However, most of them make the second one only after the first incident.

The incident in my case was educational, in the expensive sense. Early in my career I wrote a “quick script” to check a few hundred URLs across a partner company’s site. It had no gate and no per-host limit, just a thread pool and a deadline.

The partner’s staging infrastructure, which shared a load balancer with their production site, fell over. Later, the postmortem traced it to my two hundred concurrent requests. What saved the relationship was that the tool had been visible and easy to fix. One Semaphore-per-host later, plus an apology, the rerun was polite enough that their logs barely noticed it. I have kept one lesson since and built it into this article’s brief. Politeness is not courtesy; it is correctness. In other words, a concurrent tool without a gate is a bug that points at other people. Every fetcher I have written since starts with the gate, and the pool comes second.

Your Build Checklist

  1. Form 1 runs: every URL gets fetched, and the progress counter reaches total. Also, the report matches Part 4’s format, and the process exits on its own.
  2. The gate works: build a URL list with 5 URLs from one host, and watch the report timing show them paced. Then confirm no more than 2 concurrent requests in the code path.
  3. computeIfAbsent holds: with 50 tasks starting simultaneously across 10 hosts, exactly 10 semaphores exist, visible by logging gate creation.
  4. The failure path still collects. The dead-host URL from Part 4’s sample list appears as Failed with its retry. It never shows up as a crash or a hang.
  5. Form 2 is a one-line diff: diff the two files, and confirm the only change is the executor factory.
  6. Timing recorded: sequential, Form 1, and Form 2 wall times written down for the same URL list, with the sum-to-maximum shape visible.
  7. Interruptibility: Ctrl+C during a run stops the tool promptly, because every blocking call in the pipeline honors interruption.
  8. Optional stretch: exit code 1 on any failure makes the tool CI-ready. You can also add a structured-concurrency form behind –enable-preview for the all-or-nothing variant.

When NOT to Use This

  • Do not run the ungated version against any host you do not own. Also, do not treat the gate as optional polish, because it is the difference between a tool and an incident.
  • Do not point hundreds of virtual threads at a small personal site or a fragile host even with the gate. The 2-per-host number is a starting default, and some hosts deserve 1.
  • Do not use the concurrent version for a five-URL list. The sequential Part 4 tool is simpler, faster to start, and easier to debug. After all, the workload should earn concurrency.
  • Do not treat the measured numbers as universal. They are per network, per host list, per moment, and the honest comparison is same-list, same-machine, back to back.

Common Mistakes

  • Forgetting the gate’s release in finally. One exception path leaks a permit, and that host’s concurrency silently drops by one, forever. Eventually, requests queue and the tool hangs.
  • Using System.out from many threads for the report. Progress prints from tasks are fine, even interleaved. However, the report renders once, on the collecting thread.
  • Timing with System.currentTimeMillis across the run: nanoTime is the monotonic clock for durations, and wall-clock time can jump backwards.
  • Creating the HostGate per task instead of per run. Fifty gates of 2 permits per host is no gate at all. After all, one shared map is the entire point of computeIfAbsent.
  • Bare get() in the collection loop. invokeAll waits for completion, so it is safe here. However, the habit is wrong everywhere else, and the getQuietly shape is the durable one.
  • Comparing Form 1 with 8 threads against Form 2 on a single-host list. There, the gate is the binding constraint, so the comparison measures nothing. The list needs host spread to be fair.
  • Measuring once. Network variance is real, so run each version three times and compare the middle result. The profiling article will formalize the same discipline.

Key Takeaways

  • The Part 4 domain survived the concurrency upgrade untouched: sealed results, retry policy, and report are identical, because concurrency wrapped the boundary.
  • Form 1 is the executor pattern end to end: fixed pool, invokeAll, Future collection with ExecutionException unwrapping, and close-driven shutdown.
  • Form 2 is a one-line change to virtual threads. The pool-size decision disappears, and the semantically correct bound, per-host politeness, becomes the only limit.
  • The per-host gate is computeIfAbsent plus Semaphore: one atomic map entry per hostname, acquire and release in try-finally, politeness by construction.
  • Progress is an AtomicInteger incremented at task end: the threads article’s single-variable rule in its natural habitat.
  • Concurrency turns a sum into a maximum. Sequential time is the sum of fetches, while concurrent time approaches the slowest fetch, bounded by politeness.
  • Measure, then choose: same list, same machine, three runs each, and the numbers are the argument.
  • Politeness is correctness: a concurrent fetcher without a gate is a bug aimed at other people’s infrastructure.

FAQ

How do I make my Java downloader concurrent?

Submit each fetch as a Callable to a bounded ExecutorService and collect with invokeAll. Alternatively, use a virtual-thread-per-task executor for one cheap thread per task. Keep the domain logic, result types, and report unchanged, and let concurrency wrap the outside of the pipeline.

How do I limit requests per host in Java?

Keep a ConcurrentHashMap from hostname to Semaphore, filled with computeIfAbsent so exactly one gate exists per host. Then wrap each fetch in acquire before and release in finally. The map guarantees one gate per host, and the semaphore paces that host’s requests.

Should I use a thread pool or virtual threads for a downloader?

At tens to hundreds of URLs, both perform identically at human timescales. So measure both on the same list if you want evidence. Virtual threads remove the pool-size decision and scale to thousands of concurrent waits, which makes them the better default on Java 21.

How do I show progress from concurrent tasks in Java?

Share an AtomicInteger across tasks and increment it at the end of each task, after the work and the gate release. Reads are always current, increments lose nothing, and the counter is the threads article’s single-variable rule doing its job.

How do I measure concurrency speedup in Java?

Wrap each run in System.nanoTime before and after. Then run every version three times on the same URL list and machine, and compare the middle results. Expect concurrent time to approach the slowest single fetch, bounded by the per-host politeness gate.

Conclusion

Part 5 is complete, and it closed the way it opened: with your hands on the keyboard. The concurrent downloader is the same tool you built in Part 4. You promoted it through every layer of this part: pool, gate, counter. Then you ported it to virtual threads in a single-line diff, with measurements to keep the choice honest. The sealed results, the retry policy, and the report never changed, because the concurrency was always at the boundary.

The next part turns inward: the JVM itself. The next three articles cover what the virtual machine actually does when it runs your code. They show how class loading and the memory areas shape what you observed in this part. They also explain how garbage collection decides when your objects die. Finally, they teach how to profile and monitor a running application, the skills that turn “it feels slow” into a diagnosis.

Keep both forms of the concurrent downloader. Part 6 will ask you to explain what the machine was doing underneath them.

Last updated on 20 September 2026.

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *