Java

Practice: Build a Small Network Client or File Tool

Executive Summary

The tool is LinkCheck, a link checker in Java. It reads a text file of URLs, skips blanks and comment lines, and validates each line into a URI. Then it fetches each valid URI with a shared HttpClient configured with a 5-second connect timeout and 10-second request timeouts. Next, it classifies each outcome into a sealed Result hierarchy with four record cases: Fetched, Redirected, Failed, Invalid. Finally, it retries each Failed once and renders a report by pattern-switching over the results into a text block written with Files.writeString.

The design decisions are the learning. Sealed plus records makes the result space explicit and the report exhaustive by compiler check. Also, one shared HttpClient owns connection reuse, and BodyHandlers.discarding keeps the tool from downloading bodies it does not need. Meanwhile, a sequential loop is the correct concurrency model today, with the concurrent version arriving in Part 5’s practice. The checklist at the end verifies each behavior, including a deliberately dead URL, so you see the failure path work rather than assume it.

Requirements first, because practice without a spec is typing. The tool must:

Requirement Detail Exercises
Read a URL list A text file, one URL per line, blank lines and lines starting with # skipped Files.readAllLines, Path (the NIO article)
Fetch each URL HTTP GET via HttpClient, 5-second connect timeout, 10-second request timeout The networking article’s client
Classify every outcome Fetched (2xx), Redirected (3xx with location), Failed (error or bad status), Invalid (malformed line) Sealed interface plus records
Retry failures once One retry per Failed result, then accept the verdict Control flow, instanceof patterns
Write a report Counts by category plus the failure list, written to report.txt and echoed to the console Pattern switch, text blocks, Files.writeString

Design First, Code Second

Before any code, the moving parts and how they connect:

  urls.txt
     |
     |  Files.readAllLines
     v
  parse each line --------------------------------> Invalid (bad lines)
     |                                                (never leaves the ground)
     v
  fetch (shared HttpClient: 5s connect, 10s request)
     |
     v
  Result: Fetched | Redirected | Failed
     |
     |  fetchWithRetry: one retry when Failed
     v
  results: List<Result>
     |
     |  report(): pattern switch over sealed types + text block
     v
  report.txt  (Files.writeString)  +  console summary

Three design decisions, made now, carry the whole exercise. First, the result space is a sealed interface. Every possible outcome becomes a named type, and the compiler refuses any report that misses one. Second, there is exactly one HttpClient for the whole run. It owns connection pooling, so the networking article’s per-request-client mistake is designed out. Third, the loop is sequential on purpose. Fetching one URL at a time is slower and simpler, and Part 5’s practice article will parallelize this exact tool once you can reason about thread pools. Design for the concurrency you understand, upgrade when you understand more.

Step 1: The Domain Types

Start with what the program knows, not what it does. One file, LinkCheck.java, and the results it will collect:

import java.io.IOException;
import java.net.URI;
import java.net.http.*;
import java.nio.charset.StandardCharsets;
import java.nio.file.*;
import java.time.Duration;
import java.util.*;
import java.util.stream.*;

sealed interface Result permits Fetched, Redirected, Failed, Invalid {}

record Fetched(URI uri, int status, long millis)  implements Result {}
record Redirected(URI uri, URI location)           implements Result {}
record Failed(URI uri, String error)              implements Result {}
record Invalid(String raw)                        implements Result {}

Read the hierarchy as a contract. A Result is exactly one of four things, each carrying precisely what the report needs: the URI, the status, the timing, the error message, the raw line. Nothing else. In other words, this is the records article’s carriers and the interfaces article’s closed hierarchies doing the work enums cannot, because these cases carry different data.

URL validation is the first real logic, and it earns its own method:

static Optional<URI> parse(String raw) {
    try {
        var uri = URI.create(raw.strip());
        if (uri.getScheme() == null || uri.getHost() == null) {
            return Optional.empty();                     // "not-a-url", "ftp only"
        }
        return Optional.of(uri);
    } catch (IllegalArgumentException e) {
        return Optional.empty();                        // URI.create throws on garbage
    }
}

An Invalid result is data, not an exception. The file’s bad lines flow through the same pipeline as everything else and appear in the report. That is what you want from an operational tool, a complete inventory, not a crash on line three.

Step 2: The Fetcher

One shared client, one fetch method that never throws, per the networking article’s rules:

static final HttpClient CLIENT = HttpClient.newBuilder()
        .connectTimeout(Duration.ofSeconds(5))
        .followRedirects(HttpClient.Redirect.NEVER)    // a redirect is a finding, not an obstacle
        .build();

static Result fetch(URI uri) {
    var request = HttpRequest.newBuilder(uri)
            .timeout(Duration.ofSeconds(10))
            .header("User-Agent", "link-checker/1.0 (practice project)")
            .GET()
            .build();

    try {
        var start = System.nanoTime();
        var response = CLIENT.send(request, HttpResponse.BodyHandlers.discarding());
        var millis = (System.nanoTime() - start) / 1_000_000;

        return switch (response.statusCode()) {
            case int code when code >= 200 && code < 300
                                 -> new Fetched(uri, code, millis);
            case int code when code >= 300 && code < 400
                                 -> new Redirected(uri, response.headers()
                                              .firstValue("Location").map(URI::create).orElse(uri));
            case int code -> new Failed(uri, "unexpected status " + code);
        };
    } catch (IOException e) {
        return new Failed(uri, e.getClass().getSimpleName() + ": " + e.getMessage());
    } catch (InterruptedException e) {
        Thread.currentThread().interrupt();             // restore the flag, then report
        return new Failed(uri, "interrupted");
    }
}

Five decisions in that method, each from an earlier article. First, BodyHandlers.discarding is deliberate. A link checker needs the status, not the body, so the tool does not download megabytes to check a page exists. Second, the guarded patterns, case int code when …, are the modern features article’s pattern matching classifying by value range inside the same exhaustive switch. Third, the redirect decision, NEVER, means a 301 becomes reportable data, exactly what a link checker is for. Fourth, the InterruptedException handler restores the interrupt flag before converting to a Failed. That is the exceptions article’s rule for thread interruption, which Part 5 will make load-bearing. And the User-Agent identifies the tool honestly, which matters once your traffic reaches anyone’s logs.

The retry is small enough to see whole, and it uses an instanceof pattern to bind the failure it acts on:

static Result fetchWithRetry(URI uri) {
    var first = fetch(uri);
    if (first instanceof Failed f) {                    // test and bind, no cast
        System.err.printf("retrying %s after: %s%n", uri, f.error());
        return fetch(uri);                              // one retry, then accept the verdict
    }
    return first;
}

Step 3: The Report

This is where the sealed hierarchy pays for itself. Counting categories is a pattern switch the compiler checks for completeness, and the output is a text block in the shape you want to read:

static String report(List<Result> results) {
    int ok = 0, redirected = 0, failed = 0, invalid = 0;

    for (var result : results) {
        switch (result) {                                // exhaustive over the sealed types
            case Fetched    f -> ok++;
            case Redirected r -> redirected++;
            case Failed     f -> failed++;
            case Invalid    i -> invalid++;
        }
    }

    var failures = results.stream()
            .flatMap(r -> r instanceof Failed(URI uri, String error)
                          ? Stream.of("- %s (%s)".formatted(uri, error))
                          : Stream.<String>empty())
            .collect(Collectors.joining("\n"));

    return """
            Link check report
            =================
            fetched:     %d
            redirected:  %d
            failed:      %d
            invalid:     %d

            Failures:
            %s
            """.formatted(ok, redirected, failed, invalid, failures);
}

Notice the shape of the guarantee. If you later add a Timeout case to Result, this switch stops compiling until the report counts it. The failed case cannot be forgotten, the way an else branch in a ladder can. Meanwhile, the destructuring in the stream, Failed(URI uri, String error), is a record pattern pulling exactly the two components the failure line needs. This is what the modern features article meant by moving the missing-case bug from production to the compiler.

Step 4: The Main Loop

Assembly, with the exception discipline of a command-line tool. Usage errors fail fast with a message and an exit code, while everything else flows through the pipeline:

public class LinkCheck {
    public static void main(String[] args) throws IOException {
        if (args.length != 1) {
            System.err.println("usage: java LinkCheck <urls-file>");
            System.exit(2);                              // 2: usage error, distinct from 1
        }

        var lines = Files.readAllLines(Path.of(args[0]), StandardCharsets.UTF_8);

        List<Result> results = new ArrayList<>();
        for (var raw : lines) {
            var line = raw.strip();
            if (line.isEmpty() || line.startsWith("#")) {
                continue;                                // comments and blanks, tolerated
            }
            parse(line).ifPresentOrElse(
                    uri  -> results.add(fetchWithRetry(uri)),
                    ()    -> results.add(new Invalid(line)));
        }

        Files.writeString(Path.of("report.txt"), report(results));
        System.out.println(report(results));
    }
}

The exit code convention is worth keeping: 0 for a clean run, 2 for usage errors. You can also extend it so a run with failures exits 1, which is what makes the tool scriptable in CI later. The Files API from the NIO article carries all the file work, read and write, in two lines.

Run It

Build a URL file with a deliberate spread. For example, use two healthy URLs, one redirect, one dead host, one 404, one malformed line:

# urls.txt
https://openjdk.org/
https://www.rfc-editor.org/
http://github.com                          # 301: a redirect, reportable
https://this-domain-does-not-exist-9911.example/
https://httpbin.org/status/404
not-a-url

javac LinkCheck.java
java LinkCheck urls.txt

Then expect a run that takes a few seconds, a retry message on stderr for the dead host, and a report like:

Link check report
=================
fetched:     2
redirected:  1
failed:      2
invalid:     1

Failures:
- https://this-domain-does-not-exist-9911.example/ (UnknownHostException after retry)
- https://httpbin.org/status/404 (unexpected status 404)

Your exact numbers will vary, since that is the network, and the variance is instructive. In particular, the two failures demonstrate the two failure families the networking article drew. One is transport failure, the host that never answers, and the other is protocol failure, the host that answered with bad news. However, both flow through the same Failed type, and the report reads identically either way.

Extensions: Where to Take It Next

  • Exit codes with meaning: exit 1 when failed or invalid is non-zero, so a CI job can gate on the report. That makes this a real deployment guard.
  • Follow redirects by choice: add a –follow flag that switches followRedirects to NORMAL and merges Redirected into Fetched. That is one line of behavior, one line of reporting.
  • Save successful bodies: swap BodyHandlers.ofFile for discarding when you want a small site mirror. Then watch the disk fill, because that is a real lesson in scope.
  • Be a good citizen: add a fixed delay between requests to the same host. Also, read robots.txt before you point the tool anywhere that is not yours.
  • The Part 5 upgrade: fetch all URLs concurrently with an ExecutorService and collect the same List<Result>. The sealed types make the concurrent version’s correctness obvious, and that practice article walks it step by step.

How Real Systems Do This

Small tools like this are the quiet backbone of real engineering work. For example, every serious team runs link checks against documentation sites before releases and health checks against their own endpoints on a schedule. They also run file-driven batch jobs with exactly this read, act, classify, report shape. The HttpClient and HttpRequest APIs you used are the same ones behind production monitoring agents. Similarly, the sealed-result-with-report pattern is the core of every status page you have ever read.

My first real exposure to this class of tool was fixing a documentation site that shipped a release with six dead links. A customer found them, in a section the team had just rewritten. The tool that fixed the process permanently was a link checker almost exactly like this one: 150 lines, run in CI on every docs change, exit 1 on any failure. What struck me then, and has stayed with me since, was that the valuable part was not the fetching, which was routine. Instead, it was the sealed result types and the report contract, which meant the tool never lied by omission. Six years later that team still runs a descendant of that script. Meanwhile, the class of bug it was written for has never recurred. Tools you write yourself, small and honest, outlive projects.

Your Build Checklist

  1. Domain types compile: sealed Result with four records. Also, the report switch is exhaustive with no default.
  2. parse accepts http and https URIs with hosts, and rejects garbage into Invalid results without crashing.
  3. Comment lines and blanks are skipped, and the URL list tolerates trailing whitespace.
  4. One shared HttpClient: connect timeout 5s, request timeout 10s, redirects NOT followed.
  5. The dead-host URL appears in the report as Failed, with a retry visible on stderr first.
  6. The redirect URL appears as Redirected with its Location, not as a failure.
  7. report.txt is written and matches the console output; the run exits 0.
  8. Optional but recommended: exit 1 when the report contains failures. Then verify it with a CI-shaped invocation.

When NOT to Use This

  • Do not point the tool at sites you do not own without rate limits and robots.txt respect. After all, a link checker is one flag away from being an accidental DoS, and the network article’s timeout discipline does not make you a good citizen by itself.
  • Do not run this against URL lists in the tens of thousands sequentially. The design is correct and the concurrency upgrade is Part 5’s material, so upgrade the concurrency model rather than accepting hour-long runs.
  • Do not extend it into a crawler without deciding the scope first. Fetch, classify, report is a tool, whereas following every link on a page is a crawl, with storage, politeness, and legality questions attached.
  • Do not reuse the sequential pipeline for latency-sensitive work. It is deliberately simple, and pretending a blocking loop is a service is a design error the concurrency part exists to fix.

Common Mistakes

  • Creating the HttpClient inside the loop: per-request clients defeat connection pooling, the networking article’s first performance rule.
  • Downloading bodies with BodyHandlers.ofString or ofFile when the tool only needs status codes. Instead, discarding is the honest body handler for a link check.
  • Letting fetch throw instead of returning Failed: one dead host then kills the whole run. As a result, the report loses everything after the first bad URL.
  • Swallowing InterruptedException without restoring the flag: Part 5’s executors rely on that flag. So clearing it silently makes your future tool unkillable.
  • Adding a default to the report switch: over sealed types it hides the compiler’s exhaustiveness alarm, the exact protection the design bought.
  • Retrying more than once per failure: naive retry storms are how small tools become other people’s incidents. Instead, use one retry, then a verdict.
  • Testing only with healthy URLs: the dead host, the 404, and the malformed line are the test cases that matter. That is why the sample file ships them.

Key Takeaways

  • The read, act, classify, report shape generalizes: it is the skeleton of health checks, batch jobs, and operational tools across the industry.
  • Sealed result types plus exhaustive pattern switches make the report provably complete, because adding an outcome type forces every consumer to handle it.
  • One shared HttpClient with explicit timeouts is the client pattern from the networking article, now in your hands rather than on a page.
  • Failures are data: Invalid and Failed flow through the pipeline instead of crashing it, so the report is a complete inventory.
  • Retries are a policy, exactly one here, applied where a transient failure is plausible. Moreover, the network’s failure families decide the policy, not optimism.
  • InterruptedException means restore the flag and get out: the discipline becomes load-bearing the moment threads arrive in Part 5.
  • Sequential today, concurrent later. Designing for the concurrency you understand, then upgrading, is how this tool becomes Part 5’s practice article.

FAQ

What should I build to practice Java networking?

A link checker. It reads a URL list, fetches each with HttpClient, classifies outcomes into sealed result types, retries failures once, and writes a report. It exercises the HTTP client, timeouts, result modeling, and file I/O in one sitting-sized program.

How do I handle failures in a network client without crashing?

Make failures data. The fetch method catches IOException and InterruptedException and restores the interrupt flag for the latter. Then it returns a Failed record instead of throwing. One bad URL then costs one report line, not the whole run.

Why use a sealed interface for results in Java?

Because the report’s pattern switch is then exhaustive by compiler check. Every outcome type must be handled, so adding a new one breaks every switch that needs updating at compile time rather than in production.

Not by default. A redirect is a finding, so followRedirects(NEVER) reports 301s as Redirected results with their locations. That is what the tool is for. A follow mode can be an explicit flag.

How do I make this Java tool ready for CI?

Exit codes: 0 for a clean run, 1 when the report contains failures, 2 for usage errors. A CI job can then gate on the tool’s exit code. That is how a practice project becomes a deployment guard.

Conclusion

You built a real network client: file input, HTTP with timeouts, a sealed result space, a retry policy, and a report the compiler guarantees is complete. Part 4 is now yours in the strongest sense. Every mechanism in it has passed through your fingers as one working tool, and the checklist above is your evidence.

Part 5 begins the largest transformation in the course: concurrency. The next articles take the threads you have implicitly used all along and make them explicit. Then they replace most of their ceremony with ExecutorService, CompletableFuture, and virtual threads, before the Part 5 practice article returns to this exact link checker and parallelizes it properly.

Keep the tool. You will want it for the before-and-after.

Last updated on 29 September 2026.

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *