Start here: one slow response stalls the whole page#
TL;DRthe 30-second version
- Over HTTP/1.1 a connection handles one request at a time, so a slow response blocks everything queued behind it. That is head-of-line blocking.
- HTTP/2 puts many requests on one connection at once. Each request is a stream, and frames from different streams are interleaved on the wire.
- HTTP/2 still runs on TCP, which delivers bytes strictly in order. One lost packet stalls every stream. Head-of-line blocking is back, one layer down.
- HTTP/3 replaces TCP with QUIC, a transport over UDP that delivers each stream independently. A lost packet only stalls its own stream.
Open a web page and your browser makes dozens of requests: the HTML, then CSS, JavaScript, fonts, and images. In HTTP/1.1 a connection is strictly one at a time. The client sends a request, waits for the whole response, and only then sends the next one.
Say the first thing you request is a large hero image and the next two are tiny CSS files. The two small files can't come back until the big image finishes, even though they'd take milliseconds on their own. This is head-of-line blocking: the item at the front of the line holds up everything behind it.
HTTP/2's fix: many requests on one connection#
HTTP/2 keeps the same methods, paths, headers, and status codes. The headline change is multiplexing: many request/response exchanges over one connection at the same time. It takes two ideas.
- A stream is one request/response exchange inside the connection. One connection holds many streams open at once, each with its own id.
- A frame is a small chunk of one stream's data, tagged with the stream it belongs to. HTTP/2 slices every message into frames.
Because every frame carries its stream id, the sender can interleave frames from different streams on the wire. A few frames of the big image, then a few of a CSS file, then back to the image. The receiver uses the ids to reassemble each stream. A slow response no longer blocks the others.
Two supporting changes make this cheap. HTTP/2 is a binary protocol, so frames are compact to parse. And headers are compressed with a scheme called HPACK, so the same cookies aren't re-sent in full on every request. This is why gRPC is built on HTTP/2: many concurrent calls share one connection.
The catch: HTTP/2 still rides on TCP#
HTTP/2 solved head-of-line blocking inside HTTP. But it runs on TCP, and TCP has a head-of-line problem of its own. TCP delivers bytes in exactly the order they were sent. Your application never sees byte 1,001 until it has seen byte 1,000.
On the wire, all those interleaved frames travel as one ordered TCP byte stream, carried in packets. If one packet goes missing, TCP waits for it to be retransmitted before it hands any later bytes to the application. The lost packet only carried frames for one stream. But every other stream's frames that arrived safely are held behind the gap. All streams stall.
PredictHTTP/2 puts all your streams on one TCP connection. On a network with 2% packet loss, why might that single connection sometimes deliver a page more slowly than six separate HTTP/1.1 connections would?
Hint: A lost packet stalls one TCP connection. How many streams share that connection in each case?
Because TCP's in-order delivery makes one lost packet stall everything sharing that connection. On HTTP/2's single connection, a dropped packet holds back all streams until it's retransmitted, so the whole page pauses. With six HTTP/1.1 connections, a lost packet only blocks the one it happened on. The other five keep delivering. One TCP connection concentrates the blast radius of every packet loss.
HTTP/3's fix: replace TCP with QUIC#
You can't fix this by changing HTTP, because the blocking lives in TCP. So HTTP/3 stops using TCP. It runs on QUIC, a new transport built on top of UDP, the internet's bare, unordered datagram service. QUIC re-implements the good parts of TCP: reliability, congestion control, ordering. The difference is that QUIC understands streams.
QUIC tracks the bytes of every stream separately, so it knows which stream a lost packet's data belonged to. It delivers in order within a stream, but makes no ordering promise across streams. So a lost packet holds back only the one stream whose bytes were in it. The lost image packet stalls the image. The CSS and JS, whose data arrived fine, are handed over right away.
QUIC also identifies each connection by its own connection id, not by IP addresses and ports the way TCP does. So when your phone switches from Wi-Fi to cellular, the connection keeps going. That's connection migration, one reason video services favor HTTP/3 on mobile.
If this comes up in an interview#
Isn't multiplexing just HTTP/1.1 pipelining, which already existed?
No. Pipelining let a client send several requests without waiting, but the server still had to return the responses in order, so a slow first response blocked the rest. HTTP/2 interleaves frames from different streams, so responses can complete in any order.
If HTTP/2 solved head-of-line blocking, why does the problem keep coming up?
Because head-of-line blocking happens at whatever layer forces in-order handling. HTTP/2 fixed it at the application layer. TCP still delivers all bytes in order, so a lost packet stalls every stream at the transport layer. HTTP/3 was needed to fix that one.
Why UDP? Isn't UDP unreliable?
UDP itself is unreliable. QUIC builds reliability, ordering, and congestion control on top of it, per stream. QUIC couldn't change TCP, because TCP is baked into every OS kernel and every firewall on the path. UDP already passes through all of them.
Does HTTP/2 require HTTPS?
No major browser implements plaintext HTTP/2, so in practice HTTP/2 and HTTP/3 mean HTTPS. The version is negotiated during the TLS handshake via a field called ALPN.
We serve a mobile API on flaky networks. Move from HTTP/2 to HTTP/3, and what's the risk?
Lean yes. Flaky mobile hits exactly what HTTP/3 fixes: one lost packet stalling every stream, and slow setup on high latency. On a clean wired link the two perform about the same. The risks: some networks block UDP, so keep HTTP/2 as a fallback (browsers do this via the Alt-Svc header), and it costs more server CPU.
What each version buys and costs
| HTTP/1.1 | HTTP/2 | HTTP/3 | |
|---|---|---|---|
| Requests per connection | One at a time | Many (multiplexed) | Many (multiplexed) |
| Transport | TCP | TCP | QUIC (over UDP) |
| App-layer head-of-line | Yes — one response blocks the next | Solved | Solved |
| Transport head-of-line | n/a (one stream) | Yes — a lost packet stalls all streams | Solved — per-stream delivery |
| Setup round trips (with TLS) | TCP + TLS (≈2–3 RTT) | TCP + TLS (≈2–3 RTT) | QUIC = TLS folded in (≈1 RTT, 0 on resume) |
| Version | Buys you | Costs you |
|---|---|---|
| HTTP/2 | Multiplexing, header compression, one connection | Still TCP head-of-line; server push (the server sending resources you never asked for) was a flop and is being removed |
| HTTP/3 | No transport head-of-line, faster setup, connection migration | UDP is sometimes throttled or blocked by networks; more CPU (encryption in user space); newer, less battle-tested tooling |
The deeper trade is where the work moves. TCP lives in the operating system kernel and is decades-hardened. QUIC runs in user space, inside the browser or server process. That let it ship without waiting for every OS to update, but it costs more CPU per byte.
References
- RFC 9113 — HTTP/2 — The current HTTP/2 spec: streams, frames, multiplexing, HPACK.
- RFC 9000 — QUIC — The QUIC transport: per-stream delivery, connection migration, integrated TLS 1.3.
- Cloudflare — HTTP/3 vs HTTP/2 — Readable walkthrough of the head-of-line story and QUIC.
- High Performance Browser Networking — HTTP/2 — Ilya Grigorik's free chapter on streams, frames, and flow control.