Start here: why not a thread per connection?#
TL;DRthe 30-second version
- A server must hold tens of thousands of connections open at once, most of them idle. One OS thread per connection is too expensive for that. This is the C10k problem.
- The event loop is one thread in a tight cycle: ask the OS which sockets are ready, run the handler for each ready socket, repeat.
- Work per cycle scales with the number of ready sockets, not the number of open ones. An idle socket costs a few bytes of kernel state, not a thread stack.
- The catch: there is only one thread. A slow or CPU-heavy handler stalls every other connection. Heavy work goes to a worker pool.
The classic Apache-style model gives each connection its own thread. It's easy to reason about. Each connection gets its own call stack, and you write plain top-to-bottom code: read the request, do the work, write the response. But threads aren't free. Each one reserves its own stack, commonly 1 to 8 MB, plus a slot in the kernel scheduler. Run ten thousand of them and you've reserved gigabytes for stacks alone.
The deeper cost is context switching. When a thread waits on a slow socket, the OS parks it and switches to another. That means saving and restoring registers and throwing away what was in the CPU cache. Tens of thousands of switches a second is pure overhead. Dan Kegel named this wall in 1999 as the C10k problem: can one server handle ten thousand concurrent clients?
The mechanism: readiness, callbacks, repeat#
The trick has two halves. First, every socket is put in non-blocking mode. A read() with no data waiting returns at once with an error code (EWOULDBLOCK) instead of parking the thread. Second, a readiness API lets the thread ask the OS, in one call, which of its thousands of sockets have something to do right now.
That API is epoll on Linux, kqueue on BSD and macOS, and IOCP on Windows. The older select and poll calls work everywhere but don't scale, because they rescan the whole socket list on every call. With epoll you register interest in each socket once, and each wait returns only the sockets that became ready. The thread sleeps in that wait, costing no CPU, until something is ready. Then it runs the handler for each ready socket and goes back to sleep.
This shape has a name: the reactor pattern (Schmidt). One call waits on many handles, and a dispatcher routes each event to the handler registered for it. Node.js and libuv, nginx, Redis, Netty, and Python's asyncio are all this pattern with different ergonomics on top.
There's a payoff beyond saving threads. Only one thread ever touches application state, so that state needs no locks. Take Redis. A client writes SET foo bar into its socket. The command sits there until the loop notices the socket is readable, reads it, and runs it against the shared store. If two clients each have a command pending, the loop runs them one at a time. That's why Redis commands are atomic: there is never a second command running. The flip side is that one slow command, like KEYS * on a big dataset, holds the thread and blocks every other client until it finishes.
PredictYour Node service is at 5% CPU but request latency suddenly spikes to seconds for all clients at once. What's the most likely cause?
Hint: Low CPU rules out 'too much work overall.' What can one request do to all the others on a single thread?
One handler is blocking the event loop: a synchronous CPU-bound call (JSON.parse of a huge body, a sync hash, a tight loop, a blocking file read) or a microtask that keeps rescheduling itself. While that callback runs, the single thread can't touch any other socket, so every client's latency spikes together even though average CPU looks idle. The fix is to move the heavy work off the loop: a worker thread, or stream it in chunks.
The cost model: O(ready) per tick#
With C connections open and R ready this instant, one tick costs O(R) plus the handler work. Idle connections contribute nothing. They sit in the kernel's interest set as a few bytes of state and are never scanned. That is the precise sense in which one thread handles thousands of connections.
- Memory: a small struct and a socket buffer per connection (kilobytes), versus a 1 to 8 MB thread stack per connection in the blocking model.
- CPU while idle: about zero. The thread sleeps in epoll_wait() until the kernel wakes it.
- select and poll cost O(C) per call, because they re-examine every socket every time. That's why they couldn't reach C10k, and why epoll and kqueue were invented.
If this comes up in an interview#
If it's single-threaded, how is it scalable?
Because a web server's bottleneck is connection count, not CPU. Almost every connection is waiting on I/O. One thread that never blocks can interleave thousands of short I/O operations, spending CPU only on work that's ready. A thread per connection hits memory and context-switch limits long before. For CPU parallelism you run more loops.
What exactly blocks the event loop?
Anything synchronous that takes real time on the loop thread: a CPU-heavy computation like hashing or a big JSON.parse, a blocking syscall like sync file I/O or a blocking DB driver, a regex with catastrophic backtracking, or a microtask that keeps rescheduling itself. Put a synchronous 80 ms bcrypt hash in a Node handler and throughput caps near 1000/80, about 12 requests a second, with every client's latency spiking. Keep handlers short and move heavy work to a thread pool or another process.
epoll vs select: why does it matter?
select and poll pass the whole descriptor set to the kernel on every call and scan all of it, so each wait is O(n), and select has a hard cap (FD_SETSIZE). epoll and kqueue register interest once and return only the descriptors that became ready. At ten thousand connections that's the difference between rescanning 10k descriptors every loop and touching the handful that are active.
If Node is single-threaded, how does it do file I/O without blocking?
libuv runs blocking operations (file system I/O, DNS lookups via getaddrinfo) on a small background thread pool, 4 threads by default. The worker blocks on the syscall, then posts the result back to the loop, which runs your callback on the main thread. Your JavaScript is still single-threaded; only the blocking syscall moved off the loop.
Does single-threaded mean Redis can only use one core?
For command execution, yes, and that's deliberate: no locks, atomic commands. You scale across cores by running multiple instances or Redis Cluster. Redis 6 added multi-threaded network I/O for parsing and replying, but the data-structure operations stay on one thread.
Event loop vs the alternatives
| Event loop | Thread per request | Thread pool | Multi-process | |
|---|---|---|---|---|
| Concurrency ceiling | Very high (10kβ1M conns/thread) | Low β bound by thread count/RAM | Medium β bound by pool size | High β NΓ a single process |
| Memory per idle conn | Tiny (a small struct) | Large (1β8 MB stack) | Large per active worker | Tiny per conn, ΓN processes |
| CPU-bound work | Bad β blocks everyone; must offload | Good β OS preempts each thread | Good β bounded parallelism | Good β true parallelism |
| Uses many cores | No (one loop = one core) | Yes | Yes | Yes |
| Shared-state locking | None needed (single thread) | Locks/mutexes required | Locks/mutexes required | No shared memory (IPC instead) |
| Code complexity | Higher β async/callbacks | Lowest β linear blocking code | Medium | Medium β IPC, coordination |
- One loop uses one core. To use all cores you run several loops: nginx forks one worker process per core, Node uses the cluster module or worker_threads. SO_REUSEPORT lets each process bind the same port, and the kernel spreads new connections across them.
- Async code is harder to write. Control flow is inverted into callbacks, promises, or async/await, and a forgotten await or an unhandled rejection is a silent bug. Blocking code reads top to bottom; event-loop code does not.
- Most real systems combine the two: an event loop for the network front, a thread or process pool behind it for heavy work, and several loops to use all cores. nginx serving the load that Apache's prefork model needed a process per connection for is the C10k argument made concrete.
How event loops fail in production
- Blocking the loop. The number one failure. Symptom: latency for all clients spikes together while CPU may look low. Measure event-loop lag: schedule a timer for T ms and check how late it actually fires. Node exposes perf_hooks.monitorEventLoopDelay; healthy services keep p99 lag in single-digit milliseconds and alert when it climbs.
- Accept backlog overflow. The kernel finishes the TCP handshake and parks the connection in the listen backlog until your app calls accept(). If the loop is stalled, that queue fills and the kernel refuses new connections before your code or your access logs ever see them. On Linux, watch for 'SYNs to LISTEN sockets dropped' in netstat -s, or tune net.core.somaxconn.
- No backpressure. If work arrives faster than the loop drains it, internal queues grow without bound until the process is killed for memory. A client that reads responses slowly does the same thing to the socket's send buffer. Pause reads, respect the writable signal, reject, or shed load.
- Microtask starvation. In Node, promise callbacks and process.nextTick drain completely after every callback, before the loop continues. That's why an awaited promise resolves before a setTimeout(β¦, 0). A promise chain that keeps scheduling more microtasks never lets timers or I/O run.
- Errors in async callbacks. An error thrown inside a callback isn't caught by a surrounding try/catch, because that stack has already unwound. Unhandled promise rejections and missing error listeners on sockets crash the process or silently drop work.
References
- Dan Kegel β The C10K problem β the 1999 essay that framed the whole question
- Douglas Schmidt β Reactor: An Object Behavioral Pattern for Demultiplexing and Dispatching Handles for Synchronous Events β the canonical reactor-pattern paper (POSA)
- libuv β Design overview β the loop, the I/O backends, and the thread pool behind Node
- Linux man pages β epoll(7) β the scalable readiness API (and edge vs level triggering)