A waiting request still holds a thread#
TL;DRthe 30-second version
- A thread is a set of register values plus a stack. The kernel decides which thread runs on which core.
- A thread that waits for data sleeps. It uses no CPU, but it keeps its memory and its slot among the server's threads.
- Switching threads costs about 1–2 µs, plus slower memory reads until the CPU's cache refills. Each Java thread also reserves 1 MB of address space for its stack.
- So servers cap their threads. With one thread per request, that cap limits throughput to threads ÷ time per request.
Most Java web servers give each request its own thread. The thread runs the request from start to finish. When the code calls the database, the thread waits for the answer, and it can't do anything else meanwhile.
Servers keep a fixed pool of threads. Tomcat's default is 200. Little's law turns that cap into a speed limit: requests in flight equal the arrival rate times the time each one takes. Each request in flight needs a thread, so at most 200 can be in flight, and 200 ÷ 0.1 s = 2,000 requests a second.
PredictThat server has 200 threads, and each request takes 100 ms, mostly waiting. You double its CPU cores. How many more requests a second can it take?
Hint: What runs out first: cores or threads?
None. The limit is 200 threads ÷ 0.1 s = 2,000 requests a second, and extra cores don't add threads. The CPU was already mostly idle.
What a thread is: registers and a stack#
A CPU core runs one stream of instructions at a time. An 8-core machine runs 8 at once, no more.
Each core has a few dozen tiny storage slots called registers. They hold the values the core is working on right now. Two of them matter here. The instruction pointer says which instruction runs next. The stack pointer marks the top of the current thread's stack.
The stack is the thread's record of what it's in the middle of. Every method call adds a frame to it. A frame holds the method's local variables, its parameters, and a return address: where to carry on when the method finishes.
The top frame is the method running now
Objects, like the request and the response, don't live on the stack. They live on the heap, memory that every thread in the program shares. The stack only holds pointers to them.
So a thread is two things: a set of register values and a stack. Save the registers, and you've saved exactly where the thread was. Load them back, and it carries on as if nothing happened.
Who decides what runs: the scheduler#
Your program doesn't touch the hardware. The kernel does. It's the core of the operating system, and it owns the network card, the disks and the timers. When your code needs one of them, it makes a system call: a request for the kernel to do it.
The kernel also decides which thread runs on which core. That part is called the scheduler. It keeps a run queue: threads that are ready and waiting for a core. Separately, everything a thread can wait on, like each socket, has its own wait queue of sleeping threads.
A timer interrupts each core every few milliseconds. That lets the scheduler pause a running thread, even one that didn't ask to stop, and hand the core to another. This is time-slicing, and it's how 1,000 threads share 8 cores.
One thread waits for data, step by step#
- The thread calls read() on the database socket. That's a system call, so the core stops running your code and runs kernel code, still for the same thread.
- The kernel checks the socket's receive buffer, the kernel memory where arriving data waits. It's empty.
- The kernel puts the thread to sleep. It adds the thread to that socket's wait queue, and saves the thread's registers in its own record of the thread.
- The scheduler loads another thread's registers onto the core, and that thread runs. Saving one thread and loading another is called a context switch.
- The database's reply arrives. The network card copies it into memory and interrupts the core. The kernel puts the data in the socket's receive buffer.
- The kernel finds our thread on the socket's wait queue and moves it to the run queue.
- When the thread gets a core, its registers are loaded back. It finishes read(), copies the data into your program, and returns.
Two records make this work. The socket's wait queue says which thread to wake. The saved registers say where that thread was.
What a switch costs#
A context switch has two costs. The first is direct: entering the kernel, saving one thread's registers, choosing the next thread, and loading its registers. Measured on Linux, that takes about 1 to 2 microseconds.
The second cost is hidden, and usually bigger. Reading from RAM takes about 100 nanoseconds. So each core keeps a cache, a small and fast copy of memory it used recently. Reading from the nearest cache takes about 1 nanosecond.
When a thread starts running, much of its stack and objects aren't in the nearest cache. Those reads fall back to slower caches or RAM, up to 100 times slower, until the cache refills.
| Cost | Size | Why it hurts |
|---|---|---|
| Direct switch | ~1–2 µs | Enter the kernel, save and load registers, pick the next thread |
| Cold cache | up to ~100 ns per read instead of ~1 ns | The new thread's data isn't cached yet |
| Thread stack | 1 MB of address space reserved per Java thread | RAM is used only as the stack grows, but it adds up across thousands of threads |
| Kernel memory | a 16 KB kernel stack per thread (x86-64 Linux) | Paid for every thread, busy or not |
So servers cap threads#
A few thousand threads is fine. Tens of thousands means a lot of memory, and when many of them wake at once, cores spend more time switching than working. So servers keep a fixed pool, and a request waits when the pool is empty.
That leaves three choices. Raise the cap, and pay in memory and switching. Use an event loop, where one thread juggles thousands of connections with callbacks, at the cost of harder code. Or make a waiting thread cost almost nothing. That's what Java 21's virtual threads do, and they're the next read.
If this comes up in an interview#
Does a blocked thread use CPU?
No. It sleeps on a wait queue until the kernel wakes it. It still holds its stack and its place in the pool.
What's the difference between a thread and a process?
A process is a running program with its own private memory. Threads run inside a process and share its memory. Each thread has its own registers and stack.
Under the hood
- Registers, concretely. x86-64 has 16 general-purpose 64-bit registers, plus the instruction pointer, a flags register, and 16 to 32 vector registers of 128 to 512 bits. ARM64 (Apple M-series, AWS Graviton) has 31 general-purpose 64-bit registers and 32 vector registers of 128 bits. Saving a thread's registers copies a few hundred bytes to about 2–3 KB.
- The kernel stack. It's separate from your thread's stack. The kernel uses it while it runs a system call on the thread's behalf. On x86-64 Linux it has been 16 KB since kernel 3.15.
- Address translation. Programs use virtual addresses, which the CPU translates into real RAM locations. It caches recent translations in the TLB (translation lookaside buffer). Threads of one process share the same translations, so switching between them keeps the TLB useful. Switching to another process makes the old translations useless: older CPUs flush them, and newer ones tag them so they can survive the switch.
- The words. Strictly, a kernel thread runs only kernel code. People also use the term for any thread the kernel schedules. Java calls a thread that's backed one-to-one by an OS thread a platform thread.
Gotchas
- Too many runnable threads. Past a point, the CPU spends a growing share of its time switching, and latency climbs.
- A thread per idle connection. Older blocking servers held one thread per open connection, even a quiet keep-alive one. Modern servers, including Tomcat's default connector, Jetty and nginx, watch idle sockets with a handful of threads and hand over a worker thread only when a request arrives.
- Deep recursion. A thread's stack has a fixed size. Recurse too deeply and Java throws StackOverflowError.
References & further reading
- Remzi & Andrea Arpaci-Dusseau — Operating Systems: Three Easy Pieces — processes, scheduling and threads, free online
- Eli Bendersky — Measuring context switching and memory overheads for Linux threads (2018) — the 1–2 µs direct switch cost
- The Linux kernel documentation — x86 kernel stacks — the per-thread kernel stack
- The java command — -Xss — the 1 MB default thread stack on Linux/x64
- Apache Tomcat — HTTP connector — maxThreads defaults to 200