Practice: Process and Analyze a CSV with Streams
Executive Summary
This practice build shows how to analyze a CSV of sales data with a SalesAnalyzer over six columns. The program writes its own sample dataset, including one corrupted row, so the run is deterministic. A Sale record carries typed, validated fields. For example, the date is parsed as a LocalDate from the java.time article, and money is kept in cents per the Part 1 types rule. Reading uses Files.lines inside try-with-resources, skipping the header, so any file size streams one line at a time.
Parsing turns each row into Optional<Sale>, present or empty. Then flatMap(Optional::stream) splits good rows from bad in one pipeline, so invalid rows are counted and reported rather than fatal. Meanwhile, the analyses are three groupingBy collectors: summingLong for revenue by region, summingInt for units by category, and a max over the grouped entries for the top category. Finally, the extension section marks the honest boundary. split(“,”) handles this controlled dataset, but quoted fields with embedded commas are where a real CSV library takes over.
The Problem Statement
For example, a sales system exports a CSV with a header row and one row per sale. The columns are date, region, category, units sold, and unit price in cents. Management wants a summary. It needs total revenue by region, total units by category, the single top revenue category, and a count of rows that failed to import. The file may be large, so it must stream; the data is human-entered, so bad rows must not stop the run.
Requirements
- Read the CSV lazily with Files.lines, inside try-with-resources, header skipped.
- Parse each row into an immutable Sale record with typed fields: LocalDate, String, String, int, long.
- Money stays in cents throughout, per the Part 1 types rule.
- Invalid rows are counted and reported at the end; the run never fails on a bad row.
- Produce revenue by region and units by category with groupingBy collectors, and the top revenue category.
- JDK only: one complete program, no libraries.
Step 1: The Dataset and the Record
The program writes its own sample file. As a result, the demo is deterministic and it puts the exact data under test in front of you. Note the last row: corrupted units, riding along on purpose:
static Path writeSample() throws IOException {
Path csv = Files.createTempFile("sales-", ".csv");
Files.writeString(csv, """
date,region,category,units,unitPriceCents
2026-01-05,north,books,3,899
2026-01-06,south,electronics,1,24900
2026-01-07,north,books,1,450
2026-01-08,south,books,2,1299
2026-01-09,north,toys,5,999
2026-01-10,east,books,notanumber,500
""");
return csv;
}
Meanwhile, the domain type is a record, exactly as Part 2 designed them: immutable, validated at construction, with the one derived behavior the reports need:
record Sale(java.time.LocalDate date, String region, String category,
int units, long unitPriceCents) {
long totalCents() {
return (long) units * unitPriceCents; // the types article's money rule
}
}
Step 2: Parsing with a Typed Failure
To analyze a CSV safely, remember that each of the five columns can fail. However, the types article’s parse methods make each failure typed: LocalDate.parse throws DateTimeParseException, Integer.parseInt throws NumberFormatException, and a short row is caught by the column count. The parser throws, and then a wrapper turns the throw into the Optional article’s two answers, present or empty:
static Sale parse(String line) {
String[] parts = line.split(",");
if (parts.length != 5) {
throw new IllegalArgumentException("expected 5 columns: " + line);
}
return new Sale(
java.time.LocalDate.parse(parts[0].strip()), // ISO date, article 30
parts[1].strip(),
parts[2].strip(),
Integer.parseInt(parts[3].strip()),
Long.parseLong(parts[4].strip()));
}
static java.util.Optional<Sale> tryParse(String line) {
try {
return java.util.Optional.of(parse(line));
} catch (RuntimeException e) {
return java.util.Optional.empty(); // bad row: reported, never fatal
}
}
Design note, because it matters at scale: this build reports the count of bad rows, which suits a summary job. By contrast, a production import that must investigate bad rows would wrap the original exception into the custom exceptions article’s ImportException. It would carry the raw line and let the boundary decide. Counting and reporting is one honest policy, while failing loudly with the row attached is the other. However, silently skipping is neither.
Step 3: Reading Lazily and Splitting Good from Bad
The read is the NIO.2 article‘s stream-that-holds-a-handle: Files.lines inside try-with-resources, header skipped, each row mapped through tryParse. Then one line does the split, and it is the build’s centerpiece composition:
try (var lines = Files.lines(csv, StandardCharsets.UTF_8)) {
var results = lines.skip(1) // drop the header row
.map(SalesAnalyzer::tryParse)
.toList(); // Optional<Sale> per row
var sales = results.stream()
.flatMap(java.util.Optional::stream) // keep the present ones
.toList();
long rejected = results.stream()
.filter(java.util.Optional::isEmpty)
.count();
}
flatMap(Optional::stream) is where the Optional article meets the streams articles. Present Optionals dissolve into their values and empty ones vanish. As a result, the pipeline splits a mixed stream into a clean one in a single step. In effect, every row is accounted for exactly once, in one bucket or the other, with no null checks anywhere.
Step 4: The Analyses
With a clean List<Sale> in memory, the reports are the Part 2 collectors article, applied directly. Revenue by region, units by category, and the top category over the grouped entries:
Map<String, Long> revenueByRegion = sales.stream()
.collect(Collectors.groupingBy(Sale::region,
Collectors.summingLong(Sale::totalCents)));
Map<String, Integer> unitsByCategory = sales.stream()
.collect(Collectors.groupingBy(Sale::category,
Collectors.summingInt(Sale::units)));
String topCategory = sales.stream()
.collect(Collectors.groupingBy(Sale::category,
Collectors.summingLong(Sale::totalCents)))
.entrySet().stream()
.max(Map.Entry.comparingByValue()) // max over grouped entries
.map(Map.Entry::getKey)
.orElse("(no sales)");
Read each as one sentence. First, group the sales by region, summing their totals. Next, group by category, summing units. Finally, group by category again and take the entry with the highest value. The Collectors documentation lists dozens of downstream collectors, but this build uses the three that cover most reporting questions in production.
The Complete Program
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.List;
import java.util.Map;
import java.util.Optional;
import java.util.stream.Collectors;
public class SalesAnalyzer {
record Sale(java.time.LocalDate date, String region, String category,
int units, long unitPriceCents) {
long totalCents() { return (long) units * unitPriceCents; }
}
public static void main(String[] args) throws IOException {
Path csv = writeSample();
List<Optional<Sale>> results;
try (var lines = Files.lines(csv, StandardCharsets.UTF_8)) {
results = lines.skip(1)
.map(SalesAnalyzer::tryParse)
.toList(); // stream closes here, on every path
}
var sales = results.stream()
.flatMap(Optional::stream)
.toList();
long rejected = results.stream()
.filter(Optional::isEmpty)
.count();
Map<String, Long> revenueByRegion = sales.stream()
.collect(Collectors.groupingBy(Sale::region,
Collectors.summingLong(Sale::totalCents)));
Map<String, Integer> unitsByCategory = sales.stream()
.collect(Collectors.groupingBy(Sale::category,
Collectors.summingInt(Sale::units)));
String topCategory = sales.stream()
.collect(Collectors.groupingBy(Sale::category,
Collectors.summingLong(Sale::totalCents)))
.entrySet().stream()
.max(Map.Entry.comparingByValue())
.map(Map.Entry::getKey)
.orElse("(no sales)");
System.out.println("rows parsed: " + sales.size());
System.out.println("rows rejected: " + rejected);
System.out.println("revenue by region (cents): " + revenueByRegion);
System.out.println("units by category: " + unitsByCategory);
System.out.println("top category by revenue: " + topCategory);
}
static Path writeSample() throws IOException {
Path csv = Files.createTempFile("sales-", ".csv");
Files.writeString(csv, """
date,region,category,units,unitPriceCents
2026-01-05,north,books,3,899
2026-01-06,south,electronics,1,24900
2026-01-07,north,books,1,450
2026-01-08,south,books,2,1299
2026-01-09,north,toys,5,999
2026-01-10,east,books,notanumber,500
""");
return csv;
}
static Sale parse(String line) {
String[] parts = line.split(",");
if (parts.length != 5) {
throw new IllegalArgumentException("expected 5 columns: " + line);
}
return new Sale(
java.time.LocalDate.parse(parts[0].strip()),
parts[1].strip(),
parts[2].strip(),
Integer.parseInt(parts[3].strip()),
Long.parseLong(parts[4].strip()));
}
static Optional<Sale> tryParse(String line) {
try {
return Optional.of(parse(line));
} catch (RuntimeException e) {
return Optional.empty();
}
}
}
rows parsed: 5
rows rejected: 1
revenue by region (cents): {north=8142, south=27498}
units by category: {books=6, electronics=1, toys=5}
top category by revenue: electronics
Verify the numbers by hand once, because trusting a pipeline you cannot check is the mistake this practice exists to prevent. For example, north is 3 by 899 plus 1 by 450 plus 5 by 999, which is 8142 cents. Meanwhile, the corrupted row is counted, not included. Finally, electronics wins on revenue because one 24900-cent sale beats every book combined. As a result, the arithmetic is checkable, the rejection is visible, and the reports are three one-sentence collectors.
The Honest Boundary: RFC 4180
This build used split(“,”) because the dataset cooperated. However, a field may contain a comma inside quotes, a quote inside a field, or a newline inside a quoted field. At that moment, the hand-rolled split produces confidently wrong columns. That is not a Java weakness. Instead, it is the RFC 4180 specification doing what it says, because a quote-aware parser is a real state machine, not a regex. The professional rule this practice teaches is simple. split is a prototype tool for controlled, quote-free CSV. In contrast, production CSV from other systems goes through a dedicated parsing library, exactly as the strings article argued when it warned about split on real formats. In short, knowing where the boundary sits is the skill, and both sides of it are honest.
Extensions to Try Yourself
- Report the rejected rows themselves, not just the count. For example, change tryParse to return a small record holding Optional<Sale> and the raw line, then group the failures by cause.
- Use the date column. For example, filter the analysis to a specific month, or add revenue by day with groupingBy(Sale::date, TreeMap::new, summingLong(…)) for sorted output.
- Write the report to a file with Files.writeString, in the same one-call style the reading used.
- Sort the regions by revenue: collect into a List of Map.Entry, sort by value descending, and print a leaderboard.
- Handle quoted fields by hand once, as an exercise in humility. Then read the RFC again and keep a parser library in your bookmarks for the day you need it.
How Real Systems Do This
This shape, read lazily, parse with typed failure, partition, group, report, is the skeleton of production ETL. Examples include nightly sales feeds, warehouse exports, log ingestion, bank statement imports. The Java details change per system, such as a message queue instead of a file or a database insert instead of System.out. However, the pipeline is identical: typed records at the boundary, bad rows accounted for, and one-sentence collectors producing the report.
In my experience, the CSV lesson that stays with teams is the rejection policy. The first version of this exact analyzer, on a real sales feed, silently dropped bad rows and reported clean totals for six months. Meanwhile, nobody knew that an upstream region rename had corrupted a tenth of the rows. In the end, the fix was two lines: count the empties, print the count. Silent skipping is the policy that produces confident wrong reports, whereas the Optional split makes the honest policies cheap.
The performance note from the streams articles applies here directly. For a large file, the collected List<Sale> can become a stream again for the analyses. For very large files, each analysis pass re-streams from the file rather than holding everything in memory. The build as written fits millions of rows comfortably, because a Sale is small and immutable. Also, the Files.lines contract keeps the reading end at one line in memory at a time.
Decision Framework
- Is the CSV under your control, quote-free, and column-stable? Hand-rolled split, as here. Does it come from another system or users? A dedicated CSV library.
- Should a bad row stop the run? Investigate and fix: throw with the raw line, per the custom exceptions article. Summarize and continue: the Optional split with a visible count. Never silently skip.
- Is the file large? Files.lines streaming, collect only what the reports need; re-stream from the file for independent analyses rather than holding everything.
- Are the reports grouping questions? groupingBy with a downstream collector, one sentence each. Is the report a leaderboard? Max and sorted entries over the grouped map.
- Is money involved? Cents as long all the way through, converted to display format at the edge, exactly as the Part 1 types article demanded.
When NOT to Use This
- Do not hand-roll CSV parsing for external systems. Quoted fields, escaped quotes, and embedded newlines are the format’s normal cases, not its edge cases. As a result, split-based code fails on all three.
- Do not collect a multi-gigabyte file into memory to sum five columns. The reading side of this build streams, so keep the analyzing side proportional by re-streaming per report.
- Do not build the whole pipeline when the question is one total. One groupingBy over one stream answers one question, and a script’s job is allowed to be a script.
Common Mistakes
- Forgetting the header: skip(1) is not decoration. Otherwise, parsing “units” as an int throws into your rejection count, deflating both numbers.
- Silently swallowing bad rows: the pipeline runs green, the totals are wrong, and nobody notices for months. Count and report, always.
- split on unvalidated input with a regex-sensitive delimiter. Commas are safe here, but the strings article’s pipe and dot lessons transfer directly to other formats.
- Parsing numbers without a policy: this build counted failures, which is right for summaries. However, imports that must act per row need the failure wrapped with the raw line, not discarded.
- Money in double: the Part 1 rule survives every pipeline, and totalCents as long is the whole compliance program.
- Keeping Files.lines open while the analysis runs “for convenience”. Instead, the try block ends when the reading ends, and the collected list carries the data onward.
Key Takeaways
- To analyze a CSV or any similar feed, the pipeline shape is universal: read lazily, parse to typed records, split good from bad, group, report. ETL, logs, and imports are all this skeleton.
- flatMap(Optional::stream) splits a mixed stream in one step: present values flow on, empty ones are countable, and no null check appears.
- Bad rows need a policy, not an accident: count and report for summaries, throw with the raw line for imports, never skip silently.
- groupingBy with a downstream collector states a report in one sentence: region with summingLong, category with summingInt, max over entries for the top.
- Files.lines streams the file and holds the handle: collect inside the try block, close on every path, re-stream for large independent analyses.
- Money stays in cents end to end, and the date column parses through java.time’s ISO contract.
- split is honest for controlled, quote-free CSV and dishonest past it: RFC 4180’s quoted fields are where a real parser library takes over.
FAQ
How do I read a CSV file in Java?
For controlled, quote-free CSV: Files.lines inside try-with-resources, skip the header, split each line, and parse the columns into a typed record. For CSV from other systems, use a dedicated parsing library, because quoting rules are a state machine, not a split.
Can split parse CSV correctly in Java?
For fields that never contain commas, quotes, or newlines, yes, and this build proves it. The moment any field can be quoted, split produces confidently wrong columns, and RFC 4180’s quoting rules require a real parser.
How do I group data from a file in Java?
Stream the file’s lines, map each to a typed record, then collect with groupingBy and a downstream collector: groupingBy(Sale::region, summingLong(Sale::totalCents)) is a complete revenue-by-region report in one expression.
What is RFC 4180?
The specification describing CSV: fields separated by commas, rows by line breaks, and any field allowed to be wrapped in quotes so it can contain commas, quotes, and newlines. Hand-rolled splits ignore the third sentence, which is where they break.
How do I validate rows when importing a CSV?
Parse each row into typed values and let the parse methods fail: NumberFormatException and DateTimeParseException name the broken column. Wrap the failure into Optional for count-and-continue policies, or into a domain exception carrying the raw line when the import must stop and be fixed.
Conclusion
Part 3 closes with everything composed. A file streams through NIO.2, rows become validated records, and Optional separates the good from the broken. Then collectors turn survivors into reports, and the whole program is shorter than the requirements. Whenever you analyze a CSV in production, this pipeline is the shape your data jobs will take, with the boundaries this build drew around quoting and rejection policy already honest.
Next, Part 4 begins the language’s reflective machinery. Annotations come first, the metadata layer that frameworks and build tools read to do their magic, followed by reflection, modules, serialization, and the modern Java features that keep completing the toolkit.
Type your rows, count your rejects, state your reports. A pipeline that can explain itself is a pipeline you can trust.
Last updated on 7 September 2026.
