Cheap threads or simple code: you used to pick one#
TL;DRthe 30-second version
- A virtual thread is a Java thread that the kernel never sees. The JVM runs many of them on a few OS threads.
- When one waits, the JVM copies its stack frames to the heap and runs another. In most cases, the OS thread doesn't block.
- It raises how many requests can wait at once. It doesn't make any request faster.
- The limit moves to your connection pools, the services you call, and CPU. So you cap load on purpose.
With one thread per request, the code is simple: each request reads top to bottom, and a stack trace shows exactly what happened. But OS threads are expensive, so servers cap them, and a capped pool caps throughput.
The usual way out was asynchronous code: one thread juggles thousands of requests by passing callbacks. It scales, but it's harder to live with. Your logic gets split into pieces. Stack traces stop pointing at the code that failed. And once one method returns a future (a placeholder for a result that arrives later), everything that calls it has to handle futures too.
Virtual threads, final in Java 21, remove the trade-off. Keep one thread per request. Make the thread cheap.
What a virtual thread is#
A virtual thread is still a java.lang.Thread. Your code, your debugger and your stack traces work as before. The difference is that the kernel doesn't know it exists.
The JVM runs virtual threads on a small set of ordinary OS threads, called carrier threads. By default there's one carrier per CPU core. A virtual thread runs by being mounted on a carrier: its stack frames sit on the carrier's stack. When it has to wait, it's unmounted, or parked: the JVM copies its frames into an object on the heap, and the carrier runs a different virtual thread.
So there are now two schedulers. The kernel still schedules OS threads, including the carriers, onto cores. The JVM's own scheduler decides which virtual thread each carrier runs.
The kernel only sees the bottom two rows
The same read, on a virtual thread#
Here's the socket read from How OS Threads Work again, this time on a virtual thread.
- The virtual thread calls read() on the database socket. Java's socket code sees it's on a virtual thread. It keeps the socket in non-blocking mode, so the read returns at once: no data yet.
- It registers the socket with the JVM's poller. The poller is a background OS thread that asks the kernel to report when any of thousands of sockets has data (with epoll on Linux). It keeps a map from each socket to the virtual thread waiting on it.
- The JVM freezes the virtual thread: it copies its frames to the heap. The thread is now parked.
- The carrier picks up the next ready virtual thread and runs it. To the kernel, the carrier never stopped, so there's no kernel context switch.
- The database's reply arrives. The kernel tells the poller that the socket is ready.
- The poller looks the socket up in its map, finds the virtual thread, and puts it on the JVM scheduler's queue.
- Any free carrier takes it and copies its frames back. read() tries again, gets the data, and returns.
It's the same idea as with OS threads, one level up. The poller's map does the job of the socket's wait queue. The frames on the heap do the job of the saved registers.
Predict100,000 requests are each waiting 1 second on a database, on a machine with 8 carrier threads. How many carriers are stuck waiting?
Hint: Where does a waiting virtual thread live?
None. All 100,000 are parked on the heap, and the 8 carriers are free to run any work that's ready. The waiting costs memory, not threads.
What it costs#
| OS thread | Virtual thread | |
|---|---|---|
| Stack | 1 MB of address space reserved, outside the heap (RAM used only as it grows) | Only the frames in use, on the heap while parked |
| Kernel memory | A 16 KB kernel stack and a record | None: the kernel doesn't know it exists |
| Switching | Kernel switch, ~1–2 µs, then a cache refill | The JVM copies frames; no kernel involved |
| How many | Thousands | Millions (heap is the limit) |
A switch still has a cost, and the new thread's data still has to be pulled into the cache. But even a generous 1 µs per switch is small. At 5,000 requests a second, with two waits each and two switches per wait (one to park, one to resume), that's 20,000 switches a second. At 1 µs each, that's 20 ms of CPU a second, about 2% of one core.
What doesn't change is how long each request takes. As the Java team puts it, virtual threads "are not faster threads… They exist to provide scale (higher throughput), not speed (lower latency)."
Where it helps, and where it doesn't#
It helps when requests spend most of their time waiting on databases and other services, and there are many of them. The Java team's rule of thumb is more than a few thousand at once, on work that isn't CPU-bound.
It doesn't help CPU-heavy work. More threads than cores can't add computing power. And the JVM scheduler doesn't time-slice: a virtual thread keeps its carrier until it waits or finishes.
Watch for pinning. In Java 21, a virtual thread that waits inside a synchronized block, or inside native code (C code called from Java, whose frames the JVM can't copy), can't be unmounted. It holds its carrier while it waits, so 8 of them can stall all 8 carriers. On Java 21, switch a synchronized that wraps I/O to ReentrantLock. Java 24 removed the synchronized case.
And the limit moves. Without a thread cap, every request that arrives gets in. The next limit is your database connection pool, the service you call, or your CPU. A thread pool is no longer your cap, so cap load on purpose. Give each dependency a semaphore (a counter that lets at most N calls in at once), put a timeout on every call, and reject fast when you're full.
If this comes up in an interview#
Are virtual threads faster?
No. Each request takes as long as before. You can have far more of them waiting at once, which raises throughput when waiting is the bottleneck.
Do I have to rewrite my code?
Usually not. It's the same blocking code. You switch the executor (Executors.newVirtualThreadPerTaskExecutor()) or flip your framework's setting.
What's pinning?
A virtual thread that can't be unmounted while it waits, so it holds its carrier. In Java 21, that's waiting inside synchronized or native code.
Worked example: when the limit moves
A service instance gets about 570 requests a second. Each request makes one 200 ms call to another service, through an 80-thread pool that isolates that dependency. It needs 570 × 0.2 = 114 calls in flight, but the pool holds only 80, so it's already dropping calls.
Switching the web server to virtual threads doesn't fix that. The pool tops out at 80 ÷ 0.2 = 400 calls a second, and the web server can't push past it, because the pool is now the cap.
The fix is to make the call on the request's own virtual thread, and cap it with a semaphore instead. Size it from normal load plus headroom: 114 in flight normally, so allow about twice that, a cap near 250.
Then test the bad day. If the dependency slows to 2 seconds, 570 × 2 = 1,140 calls would be in flight. With no cap, they pile up in memory, time out, get retried, and slow the dependency further. With the cap, 250 slots at 2 s each let about 125 calls a second through, and the other ~445 fail fast. The instance stays healthy. That cap was the job the thread pool did by accident.
Under the hood
- The JVM scheduler is a work-stealing ForkJoinPool in FIFO mode. -Djdk.virtualThreadScheduler.parallelism sets the number of carriers (default: CPU cores). jdk.virtualThreadScheduler.maxPoolSize caps the extra carriers it may add.
- Compensation. Some operations can't unmount, such as some file I/O and, in Java 21, Object.wait(). The scheduler then adds a carrier for a while, so the others keep running.
- Always daemon, fixed priority. A virtual thread never keeps the JVM alive on its own, and ignores setPriority.
- Thread dumps. jcmd <pid> Thread.dump_to_file -format=json <file> handles millions of threads.
- ThreadLocal. Caching an expensive object per thread is fine for 200 threads and wasteful for a million, so share it or create it per use. To pass context down a call, ScopedValue (a preview in Java 21) is the lighter option.
Gotchas
- Pooling virtual threads. They're cheap to create, so make one per task. Limit concurrency with a semaphore, not a pool size.
- Old libraries. Code that waits inside synchronized pins carriers on Java 21. Find it with -Djdk.tracePinnedThreads=full or the Java Flight Recorder (JFR) event jdk.VirtualThreadPinned.
- Measuring with a closed-loop load test, where each simulated user waits for a reply before sending the next request. When the server slows, the test slows its sending too, so the slow cases never show. Use a constant-rate tool such as wrk2.
References & further reading
- JEP 444: Virtual Threads — the design, the scheduler, pinning, and the 'scale, not speed' framing
- JEP 491: Synchronize Virtual Threads without Pinning — Java 24 removes the synchronized pinning case
- wrk2 — a constant-rate load generator that measures the slow cases honestly