Java

Streams API Part 2: reduce, collect, and Collectors

Executive Summary

Streams end in values through reduce and collect. reduce folds a stream into one value with an identity element and an accumulator. For example, reduce(0, Integer::sum) totals a stream. In addition, the optional third argument, the combiner, merges partial results under parallelism. Use reduce for immutable values: sums, maximums, combined strings, and prefer IntStream’s shortcuts, sum, count, average, for primitive arithmetic. collect, by contrast, is the mutable reduction. The Collectors factory gives toList, toSet, toMap with a mandatory merge function for duplicate keys, joining for strings, and the analytics family, counting, summingInt, averagingInt. groupingBy turns a stream into a map of groups, with any collector as its downstream. Similarly, partitioningBy splits by a predicate into a true and false list. In short, use reduce for one value, collect for containers, and groupingBy for reports. That triad covers nearly every pipeline ending in production.

reduce: Folding a Stream into One Value

reduce answers “what is the one value this sequence collapses into?” It takes an identity, the answer for an empty stream, and an accumulator that combines two values into one, per the reduction lesson:

// fold to a sum: 0 is the identity (empty stream sums to 0)
int total = orders.stream()
        .map(Order::amountCents)
        .reduce(0, Integer::sum);

// fold to the longest word: "" is the identity
String longest = words.stream()
        .reduce("", (a, b) -> b.length() >= a.length() ? b : a);

// three-argument form: accumulator plus a combiner for parallel merges
int totalParallel = orders.stream()
        .reduce(0,
                (acc, order) -> acc + order.amountCents(),   // how to fold
                Integer::sum);                                // how to merge partials

The identity element is a contract, not a convenience: reduce(0, Integer::sum) is correct because zero plus anything is that thing. An identity that is not a true identity produces wrong answers under parallelism. That is why the parallel article revisits this contract with teeth.

For primitive arithmetic, however, the shortcuts from the first article do the same fold with less ceremony: mapToInt(…).sum(), .count(), .average(), .max(), and .min(). Reach for reduce when the folded value is not a built-in shortcut: longest string, combined configuration, a domain value like Money.

What reduce Is Not For

reduce builds immutable values, so forcing it to build containers is a recognized anti-pattern. Wrong code first:

// WRONG: reduce building a list: copies, mutation, and confusion
List<String> all = words.stream()
        .reduce(new ArrayList<>(),
                (list, word) -> {
                    list.add(word);          // mutating the accumulator
                    return list;             // then returning it: not a fold
                },
                (a, b) -> { a.addAll(b); return a; });

// RIGHT: collect is the mutable reduction
List<String> all2 = words.stream().collect(Collectors.toList());
List<String> all3 = words.stream().toList();   // Java 16+: immutable result

The delta: reduce demands immutable combination, so container-building versions mutate the accumulator and lie about the fold. collect exists precisely for building containers, mutating an accumulator efficiently while keeping the API honest. Note the modern ending: stream().toList(), final in Java 16, is the shortest correct form when an immutable list is acceptable.

collect: The Mutable Reduction

collect takes the JDK’s standard collectors from the Collectors class, and one table covers daily use:

Collector Produces Notes
toList() / toList() shortcut List<T> Java 16 .toList() is immutable
toSet() Set<T> Duplicates collapse
toMap(k, v) Map<K, V> Duplicate keys throw
toMap(k, v, merge) Map<K, V> Merge resolves duplicates
joining(sep, prefix, suffix) String Separator and frame optional
counting() Long As a downstream or standalone
summingInt / averagingInt Sum / average Also Long and Double variants
groupingBy(classifier) Map<K, List<T>> Groups, with downstream power
partitioningBy(predicate) Map<Boolean, List<T>> Exactly two groups
import static java.util.stream.Collectors.*;

List<String> titles = books.stream().map(Book::title).collect(toList());
Set<Isbn> isbns   = books.stream().map(Book::isbn).collect(toSet());
String csv        = titles.stream().collect(joining(", ", "[", "]"));

The toMap trap deserves the wrong-and-right treatment, because it throws in production more than any other collector:

// WRONG: two books sharing a title detonate the pipeline
Map<String, Integer> stock = books.stream()
        .collect(toMap(Book::title, Book::totalCopies));
// IllegalStateException: Duplicate key "Effective Java"

// RIGHT: the third argument resolves collisions
Map<String, Integer> stock2 = books.stream()
        .collect(toMap(Book::title, Book::totalCopies, Integer::sum));

The delta is the merge function, the same role Map.merge played in the Map article. When a key appears twice, the merge decides what the map holds. Any toMap over data with unknown duplicates without a merge function is a latent IllegalStateException, but the fix is one argument.

groupingBy: The Most Productive Method in the API

groupingBy takes a classifier, a function producing the group key. It then turns the stream into a Map from key to a list of the elements that share it. That single sentence replaces one of the most common loops in all of data processing:

// groups of orders
Map<String, List<Order>> byCategory = orders.stream()
        .collect(Collectors.groupingBy(Order::category));

// groups with a downstream collector: counts per category
Map<String, Long> countsByCategory = orders.stream()
        .collect(Collectors.groupingBy(Order::category, Collectors.counting()));

// groups with a downstream sum: the classic sales report
Map<String, Integer> salesByCategory = orders.stream()
        .collect(Collectors.groupingBy(Order::category,
                Collectors.summingInt(Order::amountCents)));

// two-way split by a predicate
Map<Boolean, List<Order>> split = orders.stream()
        .collect(Collectors.partitioningBy(o -> o.amountCents() >= 10_000));

The second argument is where groupingBy earns its reputation. Any collector can run downstream of the grouping, so counts, sums, averages, and even nested groupingBy compose like pipeline steps. What the Map article wrote as a computeIfAbsent loop, groupingBy states in one call. In addition, the container, the empty-list initialization, and the append are all handled by the API.

The Complete Example: A Sales Report in One Pipeline

import java.util.List;
import java.util.Map;
import java.util.stream.Collectors;

public class CollectDemo {

    record Order(String category, int amountCents) { }

    public static void main(String[] args) {
        var orders = List.of(
                new Order("books", 4800),
                new Order("electronics", 99900),
                new Order("books", 1500),
                new Order("toys", 4900),
                new Order("electronics", 12900));

        Map<String, List<Order>> byCategory = orders.stream()
                .collect(Collectors.groupingBy(Order::category));
        byCategory.forEach((category, list) ->
                System.out.println(category + ": " + list.size() + " orders"));

        Map<String, Integer> salesByCategory = orders.stream()
                .collect(Collectors.groupingBy(Order::category,
                        Collectors.summingInt(Order::amountCents)));
        System.out.println(salesByCategory);

        var bigOrders = orders.stream()
                .collect(Collectors.partitioningBy(o -> o.amountCents() >= 10_000));
        System.out.println("big orders: " + bigOrders.get(true).size());
    }
}
books: 2 orders
electronics: 2 orders
toys: 1 orders
{books=6300, electronics=112800, toys=4900}
big orders: 2

Read the output as three idioms. Plain grouping lists the members, and grouping with summingInt reports the totals. Finally, partitioningBy makes the split a data structure instead of two loops. The CSV practice build at the end of this part composes exactly these collectors into a full analysis pipeline.

How Real Systems Do This

Group-and-aggregate is the shape of most reporting code in production Java, such as sales by category, events by type, errors by service, request counts by status. Before streams, each of those was a loop with a Map, a containsKey or two, and an increment carefully placed after the right check. By contrast, groupingBy with a downstream collector states the whole report in one expression, and the compiler checks every type in it.

In my experience, the recurring find in stream refactoring is the report loop with a subtle bug. Either the increment sits inside the wrong branch, or the map is initialized with the wrong default. A pricing report I inherited double-counted one bundle category because the increment ran before a continue that should have guarded it. The groupingBy rewrite could not express the bug, because the grouping, the filtering, and the summing are separate typed steps. As a result, the wrong ordering simply does not compile. That is the API’s real gift, not fewer lines but fewer places for bookkeeping to hide.

The Stream specification calls reduce and collect the reduction operations, immutable and mutable respectively. The names are the design lesson. When you fold into a value, the value must never be mutated. Similarly, when you build a container, the container must be owned by the collector, not by your lambda.

Choosing Between reduce and collect

Your ending Use Example
One primitive total IntStream shortcut mapToInt(…).sum()
One immutable value reduce longest string, combined Money
A list or set collect toList / toSet filtered titles
A map from elements collect toMap with merge id to object
A map of groups groupingBy, optionally downstream orders by category with sums
A string joining, or reduce for selection logic comma-separated titles

Decision Framework

  1. Is the result one value or a container? One value: reduce or a primitive shortcut. Container: collect.
  2. Is the value a plain total, count, or average? Use the IntStream shortcut, since reduce only restates what the shortcut names.
  3. Is the container keyed by something derived from each element? toMap, with a merge function whenever duplicates are possible.
  4. Is the container groups of elements? groupingBy, with a downstream collector when the group should hold a summary rather than a list.
  5. Is the split exactly two groups by a condition? partitioningBy, because it states the dichotomy the two filters would only imply.
  6. Are you combining strings with a separator? joining; reduce-based concatenation is quadratic in spirit and wrong in practice.

When NOT to Use This

  • Do not build containers with reduce. The accumulator mutation breaks the fold contract, reads as a trick, and the parallel combiner semantics are a trap. Instead, collect exists for exactly this job.
  • Do not nest collectors five levels deep in one expression. Beyond two levels, extracting named methods or records makes the report testable and the types visible.
  • Do not group tiny, fixed data for the aesthetic. Six elements grouped in a one-off block read better as the data was written; groupingBy shines when the data outgrows the screen.

Common Mistakes

  • toMap without a merge function over duplicate-bearing data: IllegalStateException, in production, on the first duplicate.
  • A non-identity identity in reduce, like seeding a maximum with 1 instead of a real identity element. Sequential results look right, but the parallel article shows exactly how they go wrong.
  • Forgetting the combiner in the three-argument reduce. The code does not compile, and the compiler error teaches the parallel data flow better than prose.
  • Assuming .toList() returns a mutable list. Java 16’s shortcut is immutable by specification, so mutations need .collect(Collectors.toList()) or a new list.
  • Mutating the list that groupingBy produced and then reusing the same collected map. The result is shared structure that quietly diverges from the source data.
  • Using groupingBy with a classifier that can return null. The classifier’s contract forbids null keys, so the NullPointerException arrives mid-pipeline.

Key Takeaways

  • reduce folds a stream into one immutable value: identity plus accumulator, with a combiner as the third argument for parallel merges.
  • collect, by contrast, is the mutable reduction: toList, toSet, toMap, joining, and the analytics family from Collectors.
  • toMap needs a merge function whenever duplicate keys are possible; without it, the first duplicate throws IllegalStateException.
  • groupingBy produces maps of groups, and any collector can run downstream: counting, summingInt, averaging, even nested grouping.
  • partitioningBy is groupingBy for exactly two groups, keyed by a predicate.
  • stream().toList(), final since Java 16, is the immutable-list ending; Collectors.toList() is the mutable one.
  • Reduce for one value, collect for containers, groupingBy for reports: the triad that ends nearly every production pipeline.

FAQ

What does reduce do in Java streams?

It folds the stream into one immutable value using an identity element and an accumulator. For example, reduce(0, Integer::sum) totals a stream of numbers. The three-argument form adds a combiner that merges partial results when the pipeline runs in parallel.

What is the difference between reduce and collect in Java?

reduce combines values immutably, one result at a time, and suits sums and selections. In contrast, collect performs a mutable reduction into a container, list, set, or map, efficiently. The Collectors factory supplies the standard endings.

What does groupingBy do in Java?

It turns a stream into a Map keyed by a classifier function, each key holding the elements that share it. A downstream collector transforms each group, so groupingBy with summingInt is a complete category-total report in one expression.

What happens if toMap finds duplicate keys?

Without a merge function, it throws IllegalStateException on the first duplicate. The three-argument toMap takes a merge function, like Integer::sum or (a, b) -> b, that resolves collisions instead.

How do I sum a list in modern Java?

For primitives, stream to an IntStream and call sum(): list.stream().mapToInt(Order::amountCents).sum(). For objects into one domain value, use reduce with a true identity element, or IntStream’s sum when the value is a plain number.

Conclusion

Pipelines now end in values. Sums come through reduce and its primitive shortcuts, containers through collect, and complete reports through groupingBy and its downstream collectors. The Streams API is functionally complete from here, so what remains is the performance frontier.

The next article covers parallel streams. It shows when parallelStream is a real win, the ordering and identity traps that make wrong pipelines silently wrong, and the measurement discipline that decides between sequential and parallel.

Fold values, collect containers, group reports. Endings are the whole pipeline’s point.

Last updated on 12 September 2026.

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *