HotShard
Networking at scale

HTTP/2 & HTTP/3

Many requests share one connection, and a lost packet still matters.

Over HTTP/1.1, one connection works on one request at a time, so a slow response stalls everything queued behind it. HTTP/2 lets many requests share one connection at once. HTTP/3 swaps out TCP itself to fix a stall that HTTP/2 left behind. This page follows that one problem, head-of-line blocking, through the three versions.

~6 min read

Start here: one slow response stalls the whole page#

TL;DRthe 30-second version
  • Over HTTP/1.1 a connection handles one request at a time, so a slow response blocks everything queued behind it. That is head-of-line blocking.
  • HTTP/2 puts many requests on one connection at once. Each request is a stream, and frames from different streams are interleaved on the wire.
  • HTTP/2 still runs on TCP, which delivers bytes strictly in order. One lost packet stalls every stream. Head-of-line blocking is back, one layer down.
  • HTTP/3 replaces TCP with QUIC, a transport over UDP that delivers each stream independently. A lost packet only stalls its own stream.

Open a web page and your browser makes dozens of requests: the HTML, then CSS, JavaScript, fonts, and images. In HTTP/1.1 a connection is strictly one at a time. The client sends a request, waits for the whole response, and only then sends the next one.

Say the first thing you request is a large hero image and the next two are tiny CSS files. The two small files can't come back until the big image finishes, even though they'd take milliseconds on their own. This is head-of-line blocking: the item at the front of the line holds up everything behind it.

The workaround that reveals the problemBrowsers cope by opening several connections to the same server, typically up to six. A page with sixty resources still queues them six at a time, and every extra connection pays its own TCP and TLS handshake. When the fix is to open more connections, the protocol itself is the thing to fix.

HTTP/2's fix: many requests on one connection#

HTTP/2 keeps the same methods, paths, headers, and status codes. The headline change is multiplexing: many request/response exchanges over one connection at the same time. It takes two ideas.

  • A stream is one request/response exchange inside the connection. One connection holds many streams open at once, each with its own id.
  • A frame is a small chunk of one stream's data, tagged with the stream it belongs to. HTTP/2 slices every message into frames.

Because every frame carries its stream id, the sender can interleave frames from different streams on the wire. A few frames of the big image, then a few of a CSS file, then back to the image. The receiver uses the ids to reassemble each stream. A slow response no longer blocks the others.

ClientbrowserServerone connection
Three requests opened as streams 1 (image), 3 (css), 5 (js).
frame [stream 1] image…
frame [stream 3] css (done)
frame [stream 5] js (done)
frame [stream 1] image…
Small files finish early — not stuck behind the image
One connection, frames from three streams interleaved

Two supporting changes make this cheap. HTTP/2 is a binary protocol, so frames are compact to parse. And headers are compressed with a scheme called HPACK, so the same cookies aren't re-sent in full on every request. This is why gRPC is built on HTTP/2: many concurrent calls share one connection.

The catch: HTTP/2 still rides on TCP#

HTTP/2 solved head-of-line blocking inside HTTP. But it runs on TCP, and TCP has a head-of-line problem of its own. TCP delivers bytes in exactly the order they were sent. Your application never sees byte 1,001 until it has seen byte 1,000.

On the wire, all those interleaved frames travel as one ordered TCP byte stream, carried in packets. If one packet goes missing, TCP waits for it to be retransmitted before it hands any later bytes to the application. The lost packet only carried frames for one stream. But every other stream's frames that arrived safely are held behind the gap. All streams stall.

PredictHTTP/2 puts all your streams on one TCP connection. On a network with 2% packet loss, why might that single connection sometimes deliver a page more slowly than six separate HTTP/1.1 connections would?

Hint: A lost packet stalls one TCP connection. How many streams share that connection in each case?

Because TCP's in-order delivery makes one lost packet stall everything sharing that connection. On HTTP/2's single connection, a dropped packet holds back all streams until it's retransmitted, so the whole page pauses. With six HTTP/1.1 connections, a lost packet only blocks the one it happened on. The other five keep delivering. One TCP connection concentrates the blast radius of every packet loss.

HTTP/3's fix: replace TCP with QUIC#

You can't fix this by changing HTTP, because the blocking lives in TCP. So HTTP/3 stops using TCP. It runs on QUIC, a new transport built on top of UDP, the internet's bare, unordered datagram service. QUIC re-implements the good parts of TCP: reliability, congestion control, ordering. The difference is that QUIC understands streams.

QUIC tracks the bytes of every stream separately, so it knows which stream a lost packet's data belonged to. It delivers in order within a stream, but makes no ordering promise across streams. So a lost packet holds back only the one stream whose bytes were in it. The lost image packet stalls the image. The CSS and JS, whose data arrived fine, are handed over right away.

Why QUIC also connects fasterThe TLS 1.3 handshake is folded into QUIC's connection setup instead of running as a separate layer on top. HTTP/2 pays a TCP handshake and then a TLS handshake. HTTP/3 does both together, often in a single round trip. A resumed connection reuses keys cached from a previous visit and sends encrypted data in its very first packet.

QUIC also identifies each connection by its own connection id, not by IP addresses and ports the way TCP does. So when your phone switches from Wi-Fi to cellular, the connection keeps going. That's connection migration, one reason video services favor HTTP/3 on mobile.

If this comes up in an interview#

The one-linerHTTP/2 multiplexes many streams on one TCP connection, which fixes request-behind-request blocking. HTTP/3 moves to QUIC over UDP so a lost packet stalls only its own stream, and folds in the TLS handshake for faster setup.
Isn't multiplexing just HTTP/1.1 pipelining, which already existed?

No. Pipelining let a client send several requests without waiting, but the server still had to return the responses in order, so a slow first response blocked the rest. HTTP/2 interleaves frames from different streams, so responses can complete in any order.

If HTTP/2 solved head-of-line blocking, why does the problem keep coming up?

Because head-of-line blocking happens at whatever layer forces in-order handling. HTTP/2 fixed it at the application layer. TCP still delivers all bytes in order, so a lost packet stalls every stream at the transport layer. HTTP/3 was needed to fix that one.

Why UDP? Isn't UDP unreliable?

UDP itself is unreliable. QUIC builds reliability, ordering, and congestion control on top of it, per stream. QUIC couldn't change TCP, because TCP is baked into every OS kernel and every firewall on the path. UDP already passes through all of them.

Does HTTP/2 require HTTPS?

No major browser implements plaintext HTTP/2, so in practice HTTP/2 and HTTP/3 mean HTTPS. The version is negotiated during the TLS handshake via a field called ALPN.

We serve a mobile API on flaky networks. Move from HTTP/2 to HTTP/3, and what's the risk?

Lean yes. Flaky mobile hits exactly what HTTP/3 fixes: one lost packet stalling every stream, and slow setup on high latency. On a clean wired link the two perform about the same. The risks: some networks block UDP, so keep HTTP/2 as a fallback (browsers do this via the Alt-Svc header), and it costs more server CPU.

What each version buys and costs
HTTP/1.1HTTP/2HTTP/3
Requests per connectionOne at a timeMany (multiplexed)Many (multiplexed)
TransportTCPTCPQUIC (over UDP)
App-layer head-of-lineYes — one response blocks the nextSolvedSolved
Transport head-of-linen/a (one stream)Yes — a lost packet stalls all streamsSolved — per-stream delivery
Setup round trips (with TLS)TCP + TLS (≈2–3 RTT)TCP + TLS (≈2–3 RTT)QUIC = TLS folded in (≈1 RTT, 0 on resume)
VersionBuys youCosts you
HTTP/2Multiplexing, header compression, one connectionStill TCP head-of-line; server push (the server sending resources you never asked for) was a flop and is being removed
HTTP/3No transport head-of-line, faster setup, connection migrationUDP is sometimes throttled or blocked by networks; more CPU (encryption in user space); newer, less battle-tested tooling

The deeper trade is where the work moves. TCP lives in the operating system kernel and is decades-hardened. QUIC runs in user space, inside the browser or server process. That let it ship without waiting for every OS to update, but it costs more CPU per byte.

References
References

Feedback on this topic →