HotShard
Absolute basics

HTTP & APIs

The shared language every web request and response is written in.

TCP gives you a reliable pipe of bytes. But bytes alone mean nothing. The server needs to know what you're asking for, and you need to know how it answered. HTTP is the agreed shape for that exchange. The client sends a method, a path, some headers, and maybe a body. The server sends back a status code, some headers, and maybe a body. Every web page, mobile app, and public API is a stream of these messages.

~6 min read

Start here: why a shared language was needed#

TL;DRthe 30-second version
  • HTTP is a request/response protocol over a TCP (or QUIC) connection. The client sends method + path + headers + optional body. The server sends status code + headers + optional body.
  • The method is the verb. GET reads, POST creates, PUT replaces, PATCH partly updates, DELETE removes. GET and HEAD are safe (read-only). GET, HEAD, PUT, DELETE are idempotent (repeating them is harmless). POST is neither.
  • The status code is the verdict. 2xx worked, 3xx look elsewhere, 4xx the client messed up, 5xx the server broke.
  • HTTP is stateless. Every request carries its own context, such as a cookie or an Authorization token. Cache-Control and ETag headers let a response be reused instead of re-sent.

TCP hands you a reliable pipe of bytes, but bytes alone are meaningless. If a browser opens a connection and just starts sending data, the server has no idea what it means. Is this a request for a page? An upload? A login? Two programs that have never met need to agree in advance on a structure for their messages. That agreement is a protocol.

HTTP is that agreement for the web. It fixes how a request is laid out and how a response is laid out. Because every client and server follows the same rules, any browser can talk to any server it has never seen before. HTTP/1.x is plain text you can read with your own eyes: 'GET /home HTTP/1.1' is just a line you could type. That readability made it easy to implement everywhere, and it's why the web became one interoperable system instead of millions of incompatible islands.

Anatomy: the message, the verbs, the verdict#

A request and a response are each a small message with three parts: a first line, some headers, and an optional body.

PartRequestResponse
First lineGET /users/42 HTTP/1.1 — method · path · versionHTTP/1.1 200 OK — version · 3-digit status · reason
HeadersHost, Accept, Authorization: Bearer …Content-Type, Cache-Control, ETag
Body(none for a GET){"id":42,"name":"Ada"}

Two properties of a method decide how you're allowed to treat it. Safe means read-only: the request must not change server state. Idempotent means repeat-safe: sending it five times leaves the same final state as sending it once.

MethodIntentSafe?Idempotent?Typical body
GETRead a resourceYesYesNone
HEADRead headers only (no body)YesYesNone
POSTCreate / submit / 'do something'NoNoYes
PUTReplace a resource wholesaleNoYesYes
PATCHPartially update a resourceNoNoYes
DELETERemove a resourceNoYesNone

POST is the one interviewers probe. A client sends POST /payments, the server charges the card, and the response is lost on the way back. The client sees only silence. It can't tell 'never happened' from 'happened, reply lost'. The fix is an idempotency key. The client generates a unique ID and sends it in a header, such as 'Idempotency-Key: 7f3a…'. The server records the key the first time it does the work. If it sees the same key again, it returns the original result instead of charging twice. This is how payment APIs like Stripe make POST safe to retry.

Every response opens with a 3-digit status code. The first digit tells you the category before you read anything else.

ClassMeaningCommon members
1xxInformational — interim, keep going100 Continue, 101 Switching Protocols
2xxSuccess — it worked200 OK, 201 Created, 204 No Content
3xxRedirect — look elsewhere301 Moved Permanently, 302 Found, 304 Not Modified
4xxClient error — you messed up400, 401, 403, 404, 429
5xxServer error — it broke500, 502, 503, 504
  • 401 Unauthorized actually means unauthenticated: 'I don't know who you are, send credentials'. 403 Forbidden means authenticated but not allowed. The names are historically backwards. 429 Too Many Requests means you're rate-limited, often with a Retry-After header.

What a request costs, and how caching cuts it#

There's no Big-O for 'send an HTTP request'. The cost that matters is round-trips and bytes over the wire. A single HTTPS request costs a TCP handshake, then a TLS handshake, then the HTTP exchange. That's why reusing one connection for many requests (keep-alive) matters so much. The biggest lever after that is caching. If the client, or a CDN in between, can reuse a previous response, you skip the request entirely, or shrink it to a tiny 'still fresh?' check.

  • Cache-Control is the caching policy. 'max-age=60' means fresh for 60 seconds. 'no-store' means never cache. 'private' means only the browser may keep it, not a shared CDN.
  • ETag is a short fingerprint the server attaches to a response body. If the body changes, the ETag changes. The client sends it back later as 'If-None-Match' to ask whether its copy is still current.
Clienthas a cacheServerowns the body
GET /img — first visit
200 OK · ETag "v7" · 2 MB body
client caches the body + its ETag
later — is it still fresh?
GET /img · If-None-Match: "v7"
ETag still matches → unchanged
304 Not Modified · no body
Conditional GET: the ETag → 304 dance
PredictA 2 MB image is cached with ETag "v7". The client revalidates with 'If-None-Match: "v7"' and the image hasn't changed. Roughly how many bytes of body come back, and what's the status?

Hint: What does 'Not Modified' mean for the body?

Zero bytes of body, and the status is 304 Not Modified. An unchanged resource costs one small round-trip of headers instead of re-sending the 2 MB. If the image had changed, you'd get a 200 OK with the full new body and a new ETag. This is how browsers and CDNs avoid re-downloading unchanged assets on every visit.

If this comes up in an interview#

The one-linerHTTP is a stateless request/response protocol over TCP. The client sends method, path, headers and an optional body; the server replies with a status code, headers and an optional body. GET is safe, PUT and DELETE are idempotent, POST is neither.
When is it safe to retry a failed request?

When the method is idempotent (GET, PUT, DELETE), because repeating it can't double-apply. POST is risky: a lost response might mean the action already succeeded, so a retry could duplicate it. Add an idempotency key the server uses to deduplicate. That's how payment APIs make POST retry-safe.

Why send the Authorization header on every request? Doesn't the server remember me?

No. HTTP is stateless, so the server remembers nothing between requests. Each request carries its own proof of identity, a token or a cookie. That's what lets any server replica handle any request, which is what makes load balancing and horizontal scaling easy. The session data itself lives in a shared store, or is signed into the token (a JWT).

What changed between HTTP/1.1, HTTP/2, and HTTP/3?

The semantics (methods, status codes, headers) stayed the same. What changed is how messages are carried. HTTP/1.1 is text and handles one request at a time per connection, so a slow response blocks everything behind it (head-of-line blocking). HTTP/2 uses binary framing and multiplexes many requests over one TCP connection, with header compression. But one lost packet still stalls every stream, because TCP delivers in order. HTTP/3 runs over QUIC on UDP, so a lost packet only stalls its own stream, and the transport and TLS handshakes are folded together to connect in fewer round-trips.

What HTTP buys, what it costs, and which API style to put on it
  • Strength: universality. One text-based format every client and server implements, so anything can talk to anything.
  • Strength: statelessness. Self-contained requests make caching, load balancing, and horizontal scaling natural. The price is that clients re-send context (cookies, tokens) every time, and sessions and carts need an explicit store behind the stateless front.
  • Cost: verbosity. Text headers repeated on every request add bytes, and HTTP/1.1 is especially chatty. HTTP/2's header compression exists to claw this back.
  • Cost: caching has to be earned. Getting Cache-Control and ETag right is subtle, and stale or wrongly cached responses are a classic source of bugs.

Raw HTTP gives you methods, paths, and status codes. An API style is a convention for how to use them. Three dominate. REST treats everything as a resource with a URL and uses the methods as verbs and the status codes as outcomes. gRPC calls a function on the server, sends Protocol Buffers (a compact binary format) over HTTP/2, and streams natively. GraphQL has one endpoint where the client sends a query naming exactly the fields it wants. Big systems mix them: a public REST or GraphQL API at the edge, and gRPC between internal services where speed and typed contracts matter most.

RESTRPC / gRPCGraphQL
TransportHTTP/1.1 or /2, many endpointsHTTP/2 (required)HTTP, one endpoint (usually POST /graphql)
PayloadUsually JSON (text)Protobuf (compact binary)JSON, shaped by the query
SchemaOptional (OpenAPI, by convention)Required (.proto), strongly typedRequired (GraphQL schema/SDL)
StreamingLimited (SSE, long-poll)First-class (client/server/bidi)Subscriptions (often over WebSocket)
CachingExcellent — native HTTP GET cachingWeak — not HTTP-cache friendlyWeak — POSTs to one URL, needs app-level
Best usePublic web APIs, simple CRUDFast internal service-to-serviceRich UIs avoiding over-/under-fetching
References & further reading
References

Feedback on this topic →