Java

Profiling and Performance Tuning Basics in Java

Executive Summary

Profiling and performance tuning follow a five-step method. First, reproduce the problem under steady state, warm-up included, so the JIT is in the picture. Next, establish a baseline with real metrics, and profile to find where time or allocation actually concentrates. Finally, change exactly one thing and re-measure against the baseline. The metric map answers which tool to reach for. Latency at the tail means GC logs and request tracing, and throughput means load testing with a fixed workload. Similarly, memory growth means the heap histogram, and CPU means a sampling profiler.

The JDK ships the kit. It includes jstat for live heap and GC rates, plus jcmd for thread dumps, class histograms, and flags. It also includes Java Flight Recorder, JFR, a low-overhead recording you can run in production with -XX:StartFlightRecording. You then inspect the recording with the jfr tool. For deep CPU and allocation work, async-profiler adds sampling and flame graphs per its project documentation. A flame graph reads in two rules. Width is time on the CPU, height is call depth, and the wide tower is the hot spot. Micro-benchmarks belong in a harness, JMH, because hand-rolled timings lose to JIT warm-up and dead-code elimination, per the JMH project. One discipline outperforms every tool. Make one change at a time, measure it against the baseline, and revert it when the numbers say nothing.

Profiling and Performance Tuning: Measure, Change, Prove

Every competent performance investigation runs the same five steps, in the same order. Skipping a step is how weeks disappear:

1. REPRODUCE   under steady state, with warm-up: the JIT is part of the system
2. BASELINE    real metrics, not feelings: p50/p99 latency, throughput,
               allocation rate, GC time - written down, with variance
3. PROFILE     find where the time or allocation concentrates: sample, do not guess
4. CHANGE      exactly ONE thing: the fix the profile pointed at
5. PROVE       re-measure on the same workload: better, worse, or noise?
               revert what did nothing - the flag file stays a record of evidence

Two of these steps deserve their emphasis doubled. Reproduce under steady state, because a cold JVM measures its warm-up, not your change, the architecture article explained why. Also, change one thing at a time, because two simultaneous changes cancel, compound, or confound. You will not know which. As a result, the measurement of a two-change experiment is a story, not a result.

The Metric Map: Which Number Answers Which Complaint

Performance complaints come in four words, “it is slow”. So the first diagnostic act is translating the complaint into a metric. After all, each metric has its own tool and its own article:

Complaint The metric Where it shows up First tool
Requests are slow p50 and p99 latency Load test with fixed concurrent users, or production tracing JFR recording + GC log
We cannot handle the load Throughput, ops per second Load test to saturation, same workload each run jstat + a timer
Memory keeps growing Heap after GC, retention by class The histogram over time jcmd GC.class_histogram
GC pauses spike latency Pause count, duration, cadence -Xlog:gc, the GC article’s ground truth -Xlog:gc and jstat -gcutil
CPU is pegged Self-time by method A sampling profile JFR or async-profiler

The map’s discipline is the same as triage. Pick the complaint’s row and take that row’s measurement. Then let the number, not the instinct, name the next step. The GC article’s hierarchy already taught the memory rows’ fix order. Similarly, this article’s tools are how you obtain the evidence those orders demand.

The Toolkit: What Ships in the JDK

Three command families cover live triage, all introduced in the architecture article. All of them attach to a running process without restarting it:

jps -l                          // find the pid

jstat -gcutil <pid> 1000         // every second: eden/old/metaspace % used,
                                //  GC counts and total GC time: live pressure

jcmd <pid> Thread.print         // the thread dump: blocked? waiting? pool starvation?
jcmd <pid> GC.class_histogram    // what owns the heap, by class, ranked
jcmd <pid> VM.flags              // what the JVM actually started with

JFR is the recording instrument, and its defining property is overhead low enough for production. Start it on a flag or on a running process. Then inspect the file with the jfr command:

java -XX:StartFlightRecording=filename=app.jfr,duration=60s MyApp

// or on a running process:
jcmd <pid> JFR.start duration=60s filename=app.jfr

jfr summary app.jfr             // the event table: where the recording's weight sits
jfr print --events jdk.GCPause app.jfr
jfr print --events jdk.CPULoad app.jfr

For CPU and allocation hot spots, async-profiler is the standard open-source addition, one sample run, one flame graph:

asprof -d 60 -f flame.html <pid>     // sample for 60s, emit a flame graph

Here is the reading rule for any sampling output, JFR method profiles included. The samples land wherever the threads happen to be when the timer fires. Therefore, the count per method approximates its share of CPU. A method absent from the profile is not fast. Rather, it is not hot, and the distinction is the entire budget conversation.

Reading a Profile: The Flame Graph in Two Rules

A flame graph is a profile turned into a picture, and the whole skill is two reading rules:

   _______________________________ _ _ _ _ _ _
  |            main                 sample
  |   ___________     ____________
  |  |  handle  |    |  report   |
  |  | parse-JSON|   | render    |
  |__|__http-fetch|___|__________|________________
     [==== wide tower ====]      [=== narrower ===]

  RULE 1: WIDTH is TIME. The tower on the left ate most of the samples.
  RULE 2: HEIGHT is CALL DEPTH. Look at the tower's TOP: the method
          doing the work, not the callers framing it.

Beginners optimize the bottom of the tower, main, handle, because the names look important. However, the samples live at the top, in http-fetch and parse-JSON, and that is where a fix pays. Two more readings complete the skill. First, a wide, flat plateau means work spread evenly, with no single villain, so the fix is algorithmic or architectural. Second, a tower that appears under many different parents, the same method in many stacks, means one hot shared component. In that case, fixing it once pays everywhere.

Micro-Benchmarks: Why Hand-Rolled Timings Lie

“I timed it with System.nanoTime in a loop” is how performance folklore is born, and the JIT is why. A hand-rolled benchmark of a hot method faces two traps. First, warm-up means early iterations measure the interpreter. Second, dead-code elimination, the JIT removing a computation whose result is never used, means the optimized benchmark measures nothing. The professional answer is a harness that controls for both. In Java, that harness is JMH, the JDK project for exactly this problem. It offers warm-up control, forked JVMs, and protection against the optimizer’s cheating. The working rule is simple. Method-level questions get a JMH benchmark, and service-level questions get a load test against the real application. However, the middle ground, nanoTime loops in a main method, belongs to neither world and answers neither question. For a worked latency example that pairs async-profiler with JMH discipline, see Code Optimization in Java.

How Real Systems Do This

Production performance work has converged on continuous observability rather than incident-driven archaeology. JFR recordings run always or on demand with negligible overhead. GC logs ship to the metrics pipeline from day one. Meanwhile, latency dashboards watch the p99 the GC article taught you to expect spikes in. Finally, load tests gate releases with the same fixed workload, so regressions are a diff between two baselines. The engineer’s workflow is the method formalized: dashboard says slow, recording says where, one change, dashboard proves it.

The war story I tell every team is the one where the profiler beat three senior opinions. A checkout service was slow, with p99 above 400 milliseconds. The room’s theories were confident: the JSON parser, the database, the GC. All three were plausible, but only one was true. The JFR recording took ten minutes to say which. In fact, 82 percent of request time went to socket connect, the TCP handshake, before any parsing or querying happened. The cause was an infrastructure change nobody connected to latency: a per-request HttpClient created in the request path. It defeated the HttpClient’s connection pooling and paid a fresh handshake, DNS lookup included, on every call.

The fix was the networking article’s rule, one shared client, a five-line change, and p99 dropped below 90 milliseconds. The profiling was ten minutes; the arguing before it was three days. Every opinion in that room was reasonable. The recording, however, was free of the burden of reasonableness, and that is the entire argument for this article’s method.

Decision Framework

  1. Is the complaint qualitative, “slow”? Translate it to a metric first, using the map: latency, throughput, memory, GC, or CPU.
  2. Is it latency? Check the GC log before anything else, the GC article’s rule. After all, tail spikes with a clean median are GC’s signature.
  3. Is it CPU or unclear? A 60-second JFR or async-profiler sample under real load, and read the wide towers.
  4. Is it memory? Two histograms minutes apart, jcmd GC.class_histogram, and whatever grows between them is your suspect.
  5. Is the system under load in production? JFR, it is designed for exactly that, overhead-first instruments on live traffic.
  6. Is the question method-level? A JMH benchmark, never a nanoTime loop, because warm-up and dead-code elimination eat hand-rolled timings.
  7. Are you about to change something? One change, same workload, before and after, and revert anything the numbers do not justify.
  8. Is the profile empty where you expected heat? Trust it: the intuition was wrong, which is the single most common profiling outcome and the whole reason to profile.

When NOT to Use This

  • Do not optimize without a baseline. Without the “before” number, no “after” number means anything. Indeed, unmeasured optimization is indistinguishable from refactoring with extra risk.
  • Do not profile an idle dev machine and apply conclusions to production. Production has different load, different data volumes, different caches, and a different profile entirely.
  • Do not chase the visible instead of the hot. The profile’s wide tower is the budget. In contrast, a beautiful cleanup of a method the profile never mentions is unpaid work.
  • Do not run heavyweight profiling, such as heap dumps and instrumented agents, constantly on live systems. JFR’s low overhead earns continuous use, while heavier instruments earn incident use.
  • Do not micro-optimize what load testing says is fine. If the p99 is within budget, the profile is a curiosity. In that case, shipping velocity beats hypothetical milliseconds.
  • Do not benchmark cold. No warm-up means measuring the interpreter and JIT compilation. Those numbers describe the JVM’s morning, not your code.

Common Mistakes

  • Optimizing by intuition. Reasonable theories lost to ten-minute recordings in this article’s war story. Likewise, they lose the same way in every real incident room.
  • Measuring once. Variance is real at every layer. So the practice article’s three-runs-middle-result rule applies to latency and throughput as much as it did to the downloader.
  • Changing two things at once. The changes cancel, compound, or confound, and the experiment is uninterpretable. So it was not an experiment.
  • Reading the flame graph bottom-up. Width is time, and the work is at the top. Thus, optimizing the callers instead of the callee is the classic first misread.
  • Hand-rolled benchmarks: no warm-up control and no dead-code protection means the JIT optimizes the benchmark instead of measuring the code.
  • Confusing latency with throughput. They trade against each other. For example, batching raises throughput while raising latency. Tuning for the wrong one optimizes the wrong thing.
  • Skipping the GC log in any latency investigation. It is one flag, and it answers or eliminates the most common cause in one look. Besides, every other step is more expensive.
  • Keeping optimizations that did not move the number. Complexity is a cost. So revert what the baseline did not justify, and let the codebase stay as simple as the measurements allow.

Key Takeaways

  • The method is the article. Reproduce under steady state, baseline with real metrics, profile, change one thing, and prove it against the baseline.
  • The metric map translates complaints into numbers. Latency, throughput, memory, GC, and CPU each have a first tool, and the map picks it for you.
  • The JDK ships the kit: jstat for live rates, jcmd for thread dumps and histograms, and JFR for production-grade recordings. Add async-profiler for deep sampling.
  • A flame graph is two rules. Width is time and height is call depth, so the wide tower’s top is where the fix pays.
  • JFR is cheap enough to run always: continuous recordings turn performance from archaeology into a diff between two baselines.
  • Hand-rolled benchmarks lose to the JIT: warm-up and dead-code elimination eat nanoTime loops, and JMH exists to control for both.
  • Make one change at a time, measure it, and revert it when the numbers say nothing. That way, the flag and diff files stay records of evidence.
  • Trust the empty profile. When the recording contradicts the intuition, the recording wins, and that outcome is the reason profiling exists.

FAQ

How do I profile a Java application?

Reproduce the problem under steady load, and take a baseline with real metrics. Then record. Use a 60-second JFR capture, jcmd JFR.start on a running process, or an async-profiler sample for a flame graph. Read where the samples concentrate, change the one thing the profile names, and re-measure against the baseline.

What is Java Flight Recorder?

The JDK’s low-overhead recording framework. It captures method samples, GC events, allocation profiles, and VM internals while the application runs. It is light enough for production. Recordings start with -XX:StartFlightRecording or jcmd JFR.start, and the jfr tool prints or summarizes the resulting .jfr file.

What is a flame graph in Java profiling?

A visualization of sampled stacks. Width represents the number of samples, time on the CPU, and height represents call depth. The wide towers are the hot spots, and the top of a tower is the method doing the work. Meanwhile, a flat plateau means the cost is spread with no single villain.

How do I find a performance problem in a Java application?

Translate the complaint into a metric, and check the GC log first for latency complaints. Then take a 60-second recording or sample under real load, and read where the time concentrates. Change exactly one thing, re-measure on the same workload, and keep or revert based on the numbers.

What profiling tools ship with the JDK?

jps finds processes, and jstat shows live heap and GC statistics. Next, jcmd handles thread dumps, class histograms, flags, and JFR control, while jfr reads recordings. Finally, jmap and jstack take heap and thread snapshots, and javap shows bytecode. async-profiler and JMH are the standard open-source additions for sampling and micro-benchmarks.

Conclusion

Part 6 is complete. You know the machine, its memory behavior, and the instruments that watch both. Also, the method, measure, change, prove, now covers every performance conversation you will be in. The JVM articles taught you what happens when code runs. In contrast, this article on profiling and performance tuning taught you how to see it happening.

The next part changes the course’s subject from code to the machinery around code. That means build tools, test frameworks, logging, debugging, and CI. It starts with Maven, the pom.xml, and the build lifecycle that turns your source tree into deliverable artifacts. Maven is the first tool every production project requires.

Measure first. The profiler does not care how confident you are, and that is exactly its value.

Last updated on 16 September 2026.

Share this article

Leave a Reply

Your email address will not be published. Required fields are marked *