File I/O in Java: Streams, Readers, and Writers
Executive Summary
Classic file I/O in Java splits java.io into two families with one rule. InputStream and OutputStream move bytes for binary data, images, archives, and serialized blobs. In contrast, Reader and Writer move characters for text, with a charset deciding bytes to characters. The API is built as a decorator stack. For example, a raw source like FileInputStream is wrapped in a buffer, BufferedReader or BufferedInputStream, which reduces syscall counts by orders of magnitude. It is optionally wrapped in a convenience layer like PrintWriter.
Always state the charset, UTF-8 by name, even though JEP 400 made UTF-8 the JVM default in Java 18, because explicit beats inherited. try-with-resources, from the exceptions article, closes everything in reverse order on every path. However, PrintWriter is the notable trap: it swallows IOExceptions and records them in a flag you must query with checkError. Finally, read line by line for big files and all-at-once for small ones. Then let the next article’s Files class shorten the ceremonies.
File I/O in Java: Two Families, Bytes and Characters
Every class in java.io descends from one of four roots. In fact, the split is the whole design, as the essential I/O trail teaches:
BYTES (binary) CHARACTERS (text)
InputStream Reader
+-- FileInputStream +-- FileReader
+-- BufferedInputStream +-- BufferedReader (adds buffering + lines())
+-- DataInputStream +-- StringReader
OutputStream Writer
+-- FileOutputStream +-- FileWriter
+-- BufferedOutputStream +-- PrintWriter (adds print, println)
+-- PrintStream (System.out!) +-- StringWriter
Bytes are for data with no characters in it: images, archives, protocol payloads. By contrast, Readers and Writers are for text, and text is bytes plus an interpretation, which is the charset. For example, here is the wrong code first, the boundary that produced a decade of mojibake:
// WRONG: bytes decoded with an inherited assumption
try (var in = new FileInputStream("notes.txt")) {
String text = new String(in.readAllBytes()); // platform default charset:
// a gamble before JDK 18
}
// RIGHT: state the interpretation
String text = new String(bytes, StandardCharsets.UTF_8); // named, portable
The delta: historically, the default charset came from the operating system, so the same code read a file correctly on Linux and corruptly on Windows. JEP 400 fixed the default to UTF-8 from Java 18 onward, which this course builds on. However, the professional habit survives the fix. When bytes become text, name the charset, because a reader in six months should not need to know which JVM version and which operating system the code was born on. The character streams lesson makes the same argument from the other side.
The Decorator Stack: Wrap, Then Wrap Again
In practice, the API’s signature move is wrapping: each layer takes the layer below and adds one service. As a result, the stack you will write most often is a file source, a charset bridge, and a buffer:
// the classic read stack, from raw bytes to buffered lines
try (var reader = new BufferedReader(
new InputStreamReader(
new FileInputStream("notes.txt"), StandardCharsets.UTF_8))) {
String line;
while ((line = reader.readLine()) != null) { // null means end of file
process(line);
}
} // try-with-resources closes reader, which closes the layers beneath it
Read the nesting from inside out. FileInputStream produces raw bytes, and InputStreamReader decodes them into characters with a named charset. Then BufferedReader batches the decoding so the file is not hit one byte at a time, and adds readLine. Similarly, the same composition exists on the byte side, where BufferedInputStream wraps FileInputStream for the identical reason, per the byte streams lesson.
The buffering layer is not decoration. An unbuffered read hits the operating system for every byte, whereas a buffered one reads a block and serves from memory. On a 200 megabyte file, that is the difference between seconds and minutes. That is why the rule is mechanical: raw streams are wrapped, always, and the only question is which convenience layer sits on top.
Writing: OutputStream, Writer, and the PrintWriter Trap
The write side mirrors the read side: FileOutputStream for bytes, FileWriter for text, and PrintWriter for the print and println interface you have used since Part 1. Meanwhile, the append flag decides truncate versus extend:
try (var out = new PrintWriter(
new FileWriter("report.txt", StandardCharsets.UTF_8))) {
out.println("lines: " + lines);
out.println("words: " + words);
} // closed and flushed here, on every path
One trap belongs in this section because it bites exactly when correctness matters most: PrintWriter does not throw IOException. Instead, it catches it internally and sets an error flag, a design inherited from System.out. For example, here is the wrong code first:
// WRONG: a full disk writes a silently truncated file
try (var out = new PrintWriter(new FileWriter("critical.txt", StandardCharsets.UTF_8))) {
for (var line = lines(); line != null; line = next()) {
out.println(line); // if the disk fills: no exception, just a flag
}
}
// RIGHT: query the flag when the write must be certain
try (var out = new PrintWriter(new FileWriter("critical.txt", StandardCharsets.UTF_8))) {
for (var line = lines(); line != null; line = next()) {
out.println(line);
}
if (out.checkError()) { // flushes and reports accumulated failures
throw new IOException("write failed: critical.txt");
}
}
The delta: checkError() flushes and reports every swallowed failure since the last check. For logs and convenience output, the swallowing is a feature, since an interrupted stream should not kill the program. However, for data whose completeness you guarantee, the check is mandatory.
Reading Lines: The Loop and the Stream
readLine returns null at end of file. As a result, the loop is a Part 1 idiom: assign in the condition, compare with null, process in the body. And because BufferedReader.lines() returns a Stream<String>, the streams articles compose directly onto a file:
// the loop: classic, explicit, best for mutation-free processing
try (var reader = new BufferedReader(new InputStreamReader(
new FileInputStream("config.txt"), StandardCharsets.UTF_8))) {
String line;
while ((line = reader.readLine()) != null) {
if (!line.startsWith("#")) {
settings.add(line.strip());
}
}
}
// the stream: the same file as a pipeline
try (var reader = new BufferedReader(new InputStreamReader(
new FileInputStream("config.txt"), StandardCharsets.UTF_8))) {
long active = reader.lines()
.filter(line -> !line.isBlank())
.filter(line -> !line.startsWith("#"))
.count();
}
Both shapes read lazily, one line at a time. That is what keeps a multi-gigabyte log file from becoming a multi-gigabyte heap. The choice between them is the streams Part 1 rule: pipelines for filtering and transforming, loops when the body needs to mutate state you own.
The Complete Example: Write, Analyze, Report
One program exercises the full round trip. It writes a notes file, reads it back line by line, counts words with the string tools from Part 1, and writes the summary:
import java.io.BufferedReader;
import java.io.FileInputStream;
import java.io.FileWriter;
import java.io.IOException;
import java.io.InputStreamReader;
import java.io.PrintWriter;
import java.nio.charset.StandardCharsets;
public class NotesAnalyzer {
public static void main(String[] args) throws IOException {
// write a small notes file
try (var out = new PrintWriter(
new FileWriter("notes.txt", StandardCharsets.UTF_8))) {
out.println("streams move bytes");
out.println("readers move text");
out.println("buffers add speed");
}
// read it back, line by line, counting lines and words
int lines = 0;
long words = 0;
try (var reader = new BufferedReader(new InputStreamReader(
new FileInputStream("notes.txt"), StandardCharsets.UTF_8))) {
String line;
while ((line = reader.readLine()) != null) {
lines++;
for (String token : line.split("\\s+")) {
if (!token.isBlank()) {
words++;
}
}
}
}
// write the summary
try (var out = new PrintWriter(
new FileWriter("summary.txt", StandardCharsets.UTF_8))) {
out.println("lines: " + lines);
out.println("words: " + words);
if (out.checkError()) {
throw new IOException("write failed: summary.txt");
}
}
System.out.println("analyzed " + lines + " lines, " + words + " words");
}
}
analyzed 3 lines, 9 words
Every line traces to this article’s rules. Charsets are named on both writes, the read stack is wrapped three deep, and the loop ends on the null sentinel. Also, checkError guards the summary because its completeness is the program’s promise. The library practice’s CatalogStore was exactly this shape; now you can write it from memory rather than from the article.
How Real Systems Do This
In practice, file I/O in Java is the substrate under everything the course has promised so far. For example, log appenders are writers with rotation, and configuration loaders are buffered readers with a charset. Similarly, exports are print writers over file streams. Part 3’s CSV practice build reads a real dataset with this exact stack and analyzes it with the streams articles. Later, the logging article in Part 7 shows how frameworks wrap these same classes behind SLF4J.
The buffering story earned its place in a code review that has stayed with me. A batch job read a 200 megabyte export through an unbuffered stream, one byte at a time. As a result, the profile showed 99 percent of the runtime in syscalls. Wrapping the stream in a buffer was a one-line change that took the job from eleven minutes to forty seconds. Also, the code was shorter afterward. In my experience, nobody regrets wrapping a stream; everybody regrets profiling one they forgot to wrap.
Meanwhile, the charset story is the quieter cousin. The same export pipeline ran on Linux and Windows for years, and the Windows servers produced files with corrupted names. We traced the difference to the platform default charset, years before JEP 400 standardized UTF-8. The fix then and the rule now are identical: pass StandardCharsets.UTF_8 explicitly, and the operating system loses its vote.
Decision Framework
- Is the data binary or textual? Binary: InputStream and OutputStream. Text: Reader and Writer, with the charset named.
- Is the file small and simple, or large or streaming? Small: read all at once, the next article’s Files.readAllLines shortens it. Large or streaming: the buffered line loop, one line in memory at a time.
- Is the read or write on a hot path? Buffer it, always; the wrapper is one constructor call, the unbuffered cost is thousands of syscalls.
- Must the write be verified? PrintWriter plus checkError, or a plain Writer whose close can throw IOException visibly.
- Append or replace? FileWriter and FileOutputStream constructors with the append flag, decided at construction, not by habit.
- Are you copying, walking trees, checking attributes? Stop: that is the next article’s NIO.2 territory, and Files does it in fewer lines.
When NOT to Use This
- Do not read an entire large file into memory because the code is shorter. For example, readAllBytes on a two gigabyte log is an OutOfMemoryError wearing a convenience API, so the line loop exists for exactly that file.
- Do not hand-roll path handling with String concatenation and File checks. Instead, the next article’s Path and Files own that job with better errors and platform behavior.
- Do not use unbuffered raw streams in application code. The buffered wrapper is one line; the omitted wrapper is a profiler session.
Common Mistakes
- Unbuffered reads: a syscall per byte, and the program runs a hundred times slower than its data warrants.
- Relying on the platform charset: correct on one operating system, mojibake on another. JEP 400 reduced the risk and explicit UTF_8 eliminates it.
- Trusting PrintWriter to fail loudly: it swallows IOExceptions by design, so critical writes need checkError or a throwing Writer.
- Closing resources in finally instead of try-with-resources: the exceptions article’s masking bug, and the close itself throws checked IOException.
- Forgetting that readLine returns null, not an empty string, at end of file: the loop condition is the whole contract.
- Mixing the byte and character families on the same file in the same run. For example, a byte-level write between two character writes corrupts the text silently.
Key Takeaways
- Two families, one rule: byte streams for binary data and readers and writers for text. Also, name the charset explicitly as UTF-8.
- The API is a decorator stack: raw source, buffering layer, convenience layer, wrapped inside out and closed by try-with-resources in reverse.
- Buffering is mandatory on real files: one constructor call against thousands of syscalls per unbuffered read.
- readLine returns null at end of file; BufferedReader.lines() turns the same file into a Stream for the Part 3 pipelines.
- PrintWriter swallows IOExceptions by design; critical writes check checkError or use a Writer that throws.
- Read big files line by line; read small files in one call; never load gigabytes to save a loop.
- JEP 400 made UTF-8 the default charset in Java 18, and stating it explicitly survives every JVM migration.
FAQ
What is the difference between InputStream and Reader in Java?
InputStream moves bytes: binary data with no character interpretation. Reader moves characters: text decoded from bytes through a charset. Images and archives are streams; configuration and logs are readers.
Why use BufferedReader in Java?
Buffering. A raw stream reads one byte at a time through the operating system. By contrast, BufferedReader reads a block and serves lines from memory, which is typically a hundredfold speedup on real files. It also adds readLine and lines().
How do I read a file line by line in Java?
Wrap the source in BufferedReader inside try-with-resources, then loop while ((line = reader.readLine()) != null). The null return is the end-of-file signal, and the try block closes the stack on every path.
Does PrintWriter throw exceptions in Java?
No: it swallows IOExceptions and records them in an internal flag, inherited from System.out’s design. Convenience output tolerates that; data you guarantee needs checkError() after the write or a Writer whose close throws visibly.
What is try-with-resources for in file I/O?
It closes every AutoCloseable resource in reverse declaration order, on success and on failure, with suppressed exceptions preserved. For file I/O it replaces all close-in-finally code and the failure masking that came with it.
Conclusion
Classic file I/O in Java is two families and a wrapping pattern. Data is bytes or characters, buffered, closed by try-with-resources, with the charset stated and the write verified where it matters. The decorator stack generalizes, because the same layering powers network streams in Part 4 and everything the JDK ships between.
The next article shortens the ceremonies with NIO.2’s Path and Files. That modern API turns this article’s five-line stacks into one-line calls without changing a single rule underneath.
Name the charset, wrap the stream, close with the try. Files forgive none of the three.
Last updated on 16 September 2026.
