HotShard
Java memory

ZGC, and Choosing a Collector

Move objects while the program is still using them, by checking every reference as it's loaded. Then: which collector should you run?

G1 from Part 5 gets pauses down to a target you choose, but it hits a floor: every pause still copies objects with your program stopped. To go lower, a collector has to move objects while your program is still using them. This part shows how ZGC does that safely, what it costs, and how to choose between Java's collectors.

~10 min read

Moving an object while the program runs#

TL;DRthe 30-second version
  • Moving objects while the program runs is safe only if no thread ever uses an object's old address.
  • ZGC checks every reference as it's loaded (a load barrier); stale ones are fixed on the spot from a forwarding table, then written back.
  • Coloured pointers make the check one compare: a few spare bits in each reference say whether it's current.
  • Pauses stay tiny regardless of heap size; the price is throughput, GC CPU, and spare memory.
StepGC threadProgram thread
1Starts copying object O from address 1000 to 5000
2o.count = 7 writes to address 1000 (the old copy)
3Finishes the copy, points references at 5000
4Reads o.count from 5000: the 7 is lost

So moving while running only works under one strict rule: once an object starts moving, no program thread may ever use its old address. The obvious way to enforce that is to lock every object access, so the GC and the program take turns. But that means a lock on every field read and write in the whole program, which would make it run many times slower.

Check every reference as it's loaded#

A program can only touch an object through a reference, and it gets that reference by loading it from somewhere, like a field, into a local variable. So if every reference is checked at the moment it's loaded, and fixed if it's out of date, the program can never hold an old address. That's what ZGC's load barrier does:

// you wrote:  Order o = list.first;
o = load list.first
if (o is not "good")                // fast path: one compare and a branch
    o = slow_path(o, &list.first)

Almost every time, the reference is good and the program carries on. When it isn't, the slow path:

  1. Checks whether the object is in a region currently being emptied.
  2. If it has already been copied, looks up its new address in that region's forwarding table (old address → new address).
  3. If it hasn't been copied yet, the program thread copies it itself and installs the table entry. A GC thread may be copying the same object at that moment; the entry is installed atomically, so exactly one copy wins and both threads use it.
  4. Writes the corrected reference back into the field it was loaded from, so the next load takes the fast path. This is called self-healing.
  5. Returns the new address. The program never sees the old one.

Coloured pointers#

How can "is this reference good?" be one quick compare? A reference is a 64-bit number, and no heap needs all 64 bits for addresses. So ZGC uses a few spare bits in each stored reference to describe the reference itself. Those bits are its colour, hence coloured pointers.

At the start of each GC phase, ZGC changes which colour counts as good. That one change makes every reference in the heap out of date at once, without touching any of them. Each one is then checked and fixed the first time it's loaded, and it stays good after that.

PredictA field is loaded a million times during a ZGC relocation phase. How many times does its load take the slow path?

Usually once. The first load finds a stale reference, fixes it and writes the good one back into the field (self-healing). Every later load sees a good colour and takes the fast path.

What ZGC achieves, and what it costs#

ZGC marks and moves objects while the program runs. The pauses that remain are tiny, typically well under a millisecond, and they don't grow with heap size or live data; it supports heaps up to 16 TB. It also scans thread stacks concurrently, which is how it shrinks the pauses G1 keeps (Part 4). JDK 21 added a generational mode, with Part 3's young/old split; it became the default in JDK 23 and the only mode in JDK 24.

  • Throughput: every reference load runs a check, and programs load references far more often than they store them. Generational ZGC adds store barriers too, which help with marking and track old-to-young references.
  • CPU: GC threads work while your program runs, using CPU it could have had.
  • Memory: ZGC can't use compressed 4-byte references, so the same objects take more space. And the program keeps allocating while ZGC collects. If free memory runs out first, the allocating thread waits. That's an allocation stall, and to whoever is waiting it feels exactly like a pause.

Shenandoah, a collector originally built by Red Hat and included in many OpenJDK builds, solves the same problem in a similar way: barriers on references plus forwarding addresses.

The four collectors, and choosing one#

CollectorFlagMarkingMoving objectsUse it for
Serial-XX:+UseSerialGCprogram stopped, 1 GC threadprogram stoppedsmall heaps and single-CPU machines (picked automatically there)
Parallel-XX:+UseParallelGCprogram stopped, many GC threadsprogram stoppedbatch jobs: best throughput, long pauses acceptable
G1-XX:+UseG1GCalongside the program (Part 4)program stopped, a few regions per pause (Part 5)general purpose; the default on most machines
ZGC-XX:+UseZGCalongside the programalongside the programtiny pauses on large heaps, with spare CPU and memory

Parallel just adds GC threads. Each row after that removes a stop-the-world phase and pays for it with more work while the program runs, in barriers and GC threads. So choose by which kind of pause you can't accept: Parallel when only total time matters, G1 as the default, ZGC when your longest pauses are unacceptable, especially on a large heap, and you can afford the extra CPU and memory.

And whichever you choose, the series' lessons hold. Short-lived objects are cheap; they die young and are never visited (Part 3). Large long-lived data is what costs, because every marking walks it and every compaction copies it (Parts 2, 4, 5). A bigger heap means fewer collections, not less work per collection, and with G1 and ZGC it also gives the collector room to finish before memory runs out.

If this comes up in an interview#

The one-linerZGC marks and moves objects while the program runs, so its pauses stay tiny regardless of heap size. A load barrier checks every reference as it's loaded, using coloured pointers and forwarding tables, so the program never uses an object's old address. The price is throughput, GC CPU, and spare heap to avoid allocation stalls.
Load barrier vs write barrier?

A write barrier runs extra code when a reference is stored (card marking, SATB). A load barrier runs it when a reference is read; ZGC uses one to fix stale references before the program can use them.

What is an allocation stall?

ZGC running out of free memory before a concurrent cycle finishes, so the allocating thread waits. It shows up as latency, like a pause.

G1 or ZGC?

G1 by default. ZGC when the slowest 1% of requests (p99 latency) matter more than a few percent of throughput, and the heap is large enough that G1's copying pauses hurt.

Gotchas
  • Read the GC log before tuning: -Xlog:gc*:file=gc.log shows every collection's type, cause, duration, and heap before and after.
  • Leaks are reachability bugs, not GC bugs. Take a heap dump (jcmd <pid> GC.heap_dump heap.hprof) and look at the path to GC roots of whatever keeps growing.
  • System.gc() is only a request, and it usually triggers a full collection. Avoid it.
  • A ZGC heap sized tightly to live data invites allocation stalls; give it headroom.
Under the hood
  • In generational ZGC, the colour bits are the low bits of a reference stored in the heap (the address is shifted above them); the load barrier strips them, so a loaded reference is a plain address.
  • ZGC's three remaining pauses (mark start, mark end, relocate start) do a small, fixed amount of work, like switching the good colour and syncing threads, which is why they don't grow with heap size.
  • Earlier, non-generational ZGC mapped the same physical memory at several virtual addresses, one per colour; generational ZGC removed that trick.
References & further reading
References

Feedback on this topic →