Code Optimization in Java: Minimize Latency with Profiling

Java latency playbook: allocation and I/O techniques, pooling, async-profiler and JMH guidance, with before/after discipline.

Executive Summary: Cute microbenchmarks rarely match p99 pain in production. This post starts from a string-merge micro-optimization, then covers allocation control, I/O, pooling, async-profiler, and JMH discipline for real Java services. For backend engineers chasing latency with evidence, not folklore flags.

In software development, code optimization matters most when latency shows up in user-facing p99s, not when a microbenchmark looks cute. A well-chosen allocation strategy or I/O pattern can separate a sluggish service from one that stays under budget. Here I will show a concrete instance of code optimization in Java, then expand into the techniques I reach for on real services: allocation control, I/O, pooling, and profiling with async-profiler and JMH-level discipline.

Recently, while solving a problem that involved merging two strings in alternating order, I hit an interesting performance challenge. My initial solution sat around 2ms in the judge environment. After applying a few targeted changes I saw 0ms on that same harness (timer granularity warning below). I will keep that anecdote, then show how the same instincts transfer to production Java.

The problem: merge two strings alternately

Given word1 and word2, merge by alternating characters. If one string is longer, append the remainder.

Initial solution (StringBuilder)

class Solution {
  public String mergeAlternately(String word1, String word2) {
    StringBuilder result = new StringBuilder();
    int i = 0;
    while (i < word1.length() || i < word2.length()) {
      if (i < word1.length()) {
        result.append(word1.charAt(i));
      }
      if (i < word2.length()) {
        result.append(word2.charAt(i));
      }
      i++;
    }
    return result.toString();
  }
}

Observed: about 2ms on that platform.

It is correct and readable. StringBuilder is usually the right default. On a hot path with known final size, we can still do better.

Optimized solution (pre-sized char[])

class Solution {
  public String mergeAlternately(String word1, String word2) {
    int m = word1.length();
    int n = word2.length();
    char[] result = new char[m + n];
    int i = 0, j = 0, k = 0;
    while (i < m || j < n) {
      if (i < m) {
        result[k++] = word1.charAt(i++);
      }
      if (j < n) {
        result[k++] = word2.charAt(j++);
      }
    }
    return new String(result);
  }
}

Observed: 0ms on the same harness.

What actually changed

  • No dynamic growth: StringBuilder may resize; a correctly sized char[] does not.
  • Fewer virtual calls per character: direct array stores vs repeated append.
  • Contiguous writes: friendlier to CPU caches on tight loops.
  • JIT-friendly shape: simple counted loops optimize well.

Caveat on “0ms”: online judges often use coarse timers. Treat 2ms → 0ms as “noticeably cheaper,” not as a scientific speedup factor. For real claims, use JMH with warmup, multiple forks, and report percentiles.

Good development is not only making it work. It is knowing when microseconds matter and when readability wins. (Yes, still channeling Ranchoddas Shamaldas Chanchad energy.)

Before / after mindset for production latency

Symptom Before (common) After (targeted)
High young-gen GC / allocation rate Lots of short-lived objects per request Reuse buffers, avoid autoboxing on hot paths, pre-size collections
p99 spikes under load Blocking I/O on event threads; lock contention Async or bounded thread pools; shrink critical sections; profile waits
“Fast enough” locally, slow in prod Microbench without warmup; different GC flags Match prod JVM flags; measure with async-profiler in staging
JSON/HTTP overhead dominates Reflection-heavy mappers allocating per field Reuse parsers, consider afterburner/binary codecs where justified

Technique 1: allocation and object churn

Allocation is not free. On a service handling tens of thousands of RPS, per-request junk shows up as GC pauses and CPU in Thread.allocate / TLAB refill.

  • Pre-size ArrayList, HashMap, StringBuilder when length is known.
  • Prefer primitives and primitive collections on numeric hot paths.
  • Avoid concatenating strings in loops; build once.
  • Watch hidden allocations: lambdas capturing locals, stream pipelines on tight loops, Optional in inner loops, autoboxing in maps of Integer.
  • Reuse thread-local or pooled buffers for encoding/decoding when profiling proves allocation dominates. Do not pool prematurely; pools add complexity and leaks.

Heuristic: if async-profiler CPU view shows significant time in GC or in allocator stubs, open the allocation view and fix the top sites before rewriting algorithms.

Technique 2: I/O and concurrency

  • Never block the event loop (Netty / reactive) with JDBC or file I/O. Offload to a bounded elastic pool.
  • Use connection pools (HikariCP-class) with timeouts that fail fast. A hung DB call that waits minutes is a latency amplifier for everyone else.
  • Batch where the protocol allows (multi-get, pipelining) instead of chatty sequential calls.
  • Set explicit read/connect timeouts. Infinite waits become mysterious p99s.
  • Prefer zero-copy / direct buffers only when profiling shows copies dominate; clarity first.

Technique 3: pooling (connections, buffers, workers)

Pooling is a latency tool when creation cost or admission control matters:

  • DB and HTTP client pools: size from measured concurrency, not folklore (not “always 200”).
  • Thread pools: named, bounded queues, explicit rejection policy. Unbounded queues hide overload until memory dies.
  • Object pools: rare. Use when instances are expensive (crypto engines, large direct buffers) and lifecycle is clear.

Always pair pools with metrics: active/idle, wait time to borrow, rejection count.

Technique 4: profiling with async-profiler

async-profiler is my default for production-shaped JVMs. CPU and allocation flame graphs beat guessing.

# CPU (example; match your install path and JDK)
./asprof -e cpu -d 60 -f /tmp/cpu.html <pid>

# Allocations
./asprof -e alloc -d 60 -f /tmp/alloc.html <pid>

# Wall-clock (catch blocking)
./asprof -e wall -d 60 -f /tmp/wall.html <pid>

Workflow I use:

  1. Reproduce load in staging (realistic payload sizes).
  2. Capture CPU + alloc + wall profiles under that load.
  3. Fix the tallest frames that you own (not JDK noise unless it points to your API misuse).
  4. Re-measure p50/p95/p99 and error rate. Keep a short before/after note in the PR.

For continuous insight, JFR + async-profiler integrations or vendor APM are fine; the habit matters more than the brand.

Technique 5: JMH-level measurement discipline

JMH exists so we stop lying to ourselves with System.nanoTime around a single call.

  • Warmup iterations until scores stabilize.
  • Multiple forks to dilute machine noise.
  • Avoid dead-code elimination with Blackhole consumption.
  • State scopes that match reality (@State(Scope.Thread) vs Benchmark).
  • Report percentiles if tail latency is the product concern; averages hide pain.

Microbenchmarks validate an algorithm hypothesis. They do not replace service-level SLOs. The string-merge anecdote is a teaching micro-optimization; your checkout API cares about DB, JSON, auth, and GC together.

JVM and GC knobs (only after evidence)

  • Start with a modern LTS JDK and G1 or ZGC depending on heap and pause goals; do not cargo-cult flags from 2015 blog posts.
  • Size heap from live set + headroom; too small causes thrashing, too large can stretch pauses depending on collector.
  • Log GC (-Xlog:gc*) during incidents; correlate with request latency charts.
  • Disable biased optimizations only when profiles say so.

A compact before/after playbook

  1. Before: define the SLO (for example p99 < 150ms at 2k RPS) and capture baseline profiles.
  2. Change one class of issue (allocations or lock contention or I/O), not five at once.
  3. After: same load script, same JVM flags, compare profiles + latency histogram.
  4. Document: flame graph screenshots and the one-paragraph cause in the PR for the next engineer.

When not to optimize

If the method runs once per request and allocates a few kilobytes, leave the clear StringBuilder version. Optimize when profiles or SLOs demand it. Readability and correct concurrency beat clever char[] tricks in cold code.

For broader backend context on this site, see the microservices material when latency is dominated by service boundaries rather than a single tight loop. For AI-assisted refactors, keep humans in the review loop using the practices in AI tools for developers.

Takeaways

  • Pre-size when the final length is known; measure with honest tools.
  • Attack allocations, blocking I/O, and pool misconfiguration before exotic JVM flags.
  • async-profiler + JMH cover “what is hot” and “is this micro change real.”
  • Publish before/after histograms, not vibes.

The difference between 2ms and 0ms on a toy problem is a reminder: small structural choices compound. Across a large Java estate, those choices show up as cheaper CPU bills and steadier tails. Balance that against code your teammates can still change on a Monday morning.

Related reading: Dependency Injection in Spring Boot, PostgreSQL thesis, and Java OOP principles.

Related: Understanding JVM Internals: From Source Code to Runtime

Share this article

5 thoughts on “Code Optimization in Java: Minimize Latency with Profiling”

  1. Marcus Torres

    Focused write-up on cutting Java latency without boiling the ocean. Loop and allocation notes were practical.

  2. Myra

    Optimization tips that matter beat microbenchmark theater. Latency section stayed grounded in production habits.

  3. Kabir

    Calendar block titled simply: apply the checklist. Live demo path follows java latency.

  4. Ben

    Shifted our spike scope down after reading. Reviewers understood the hot path allocation change.

  5. Arjun

    Needed this. Thanks for covering optimization tips that matter on real projects.

Leave a Reply

Your email address will not be published. Required fields are marked *