Start here: why this pattern exists#
TL;DRthe 30-second version
- Almost everything online is one pattern repeated: a client asks (a request) and a server answers (a response). One ask plus one answer is a round-trip.
- A server isn't special hardware. It's a program running a loop: wait for a request, do the work, send a reply, repeat.
- The cost that matters most is latency, how long one round-trip takes. Fewer round-trips means a faster app.
- A request can get no answer at all. The client can't tell 'slow' from 'never', and that one fact is the seed of nearly every hard distributed-systems topic.
Imagine two computers. One holds your photos, the other wants to show them. They could each try to reach into the other's memory directly, but that would be chaos: no order, no permissions, no way to know who is allowed to do what. Software needs a simple, predictable way for one program to ask another for something across a network.
The client-server pattern is that way. One side asks (the client) and the other answers (the server). The roles are fixed for a conversation: the client always speaks first, and the server only ever replies. So there is always exactly one party in charge of starting, and one party in charge of the data.
The round-trip, what it costs, and when it gets no answer#
A client is whatever starts the conversation: your browser, a mobile app, or another service. A server is a program on some machine that listens for requests and replies to them. The server never calls you out of the blue. It only responds.
- A request carries four parts. The method says what to do: GET to read, POST to create, PUT or PATCH to update, DELETE to remove. The path says which resource, like /users/42. Headers carry extra info, such as who's asking and what format they want. The optional body carries the data being sent, like a form or JSON.
- A response carries a status code (a 3-digit verdict: 200 OK, 404 Not Found, 500 Server Error), headers (format, caching hints, cookies), and an optional body (HTML, JSON, an image).
There's no Big-O for 'send a request'. The dominant cost isn't computation, it's the journey. Latency is how long one round-trip takes, from sending the request to the full response arriving. Distance dominates it, because the speed of light is a hard limit. Same-datacenter latency is often under a millisecond. Across the world it can be 100 to 300 ms before the server even starts working.
PredictA page makes 5 requests to a server 100 ms away. If it sends them one after another (each waits for the last), it takes about 500 ms. If it sends all 5 at once over one open connection, roughly how long?
Roughly 100 ms, about one round-trip, not five. The requests travel in parallel and the responses come back together, so you pay the travel time once instead of five times (assuming the server handles them concurrently). The bottleneck is the round-trip, and parallel round-trips overlap.
Now the sharp edge. The request crosses a network to reach the server, and the response crosses it back. If the server is down, or the network drops a message, the client gets silence. Did the server never receive the request? Receive it but die before replying? Reply, but the reply got lost on the way back? From the client's side all three look identical. This is called partial failure: part of the system worked, part didn't, and you can't be sure which.
- Timeout: the client can't wait forever, so it sets a deadline and declares failure when it passes. Too short and you give up on healthy-but-slow servers. Too long and a dead server hangs your whole app.
- Retry: send it again. Often that works, because the blip was temporary. But if the first request actually succeeded and only the reply was lost, retrying does the action twice.
- Idempotency: the fix for safe retries. An idempotent request can be sent any number of times with the same end result ('set balance to 100'). A non-idempotent one ('add 100 to the balance') is unsafe to repeat. Design requests to be idempotent, or tag them with a unique key the server can de-duplicate.
If this comes up in an interview#
What exactly is a server, then?
A program that runs a loop: wait for an incoming request, do some work, send back a response, repeat. The word also gets used for the machine the program runs on, but the useful definition is the program. Any computer can run one, including your laptop. 'Server' is a role, not a kind of hardware.
REST vs RPC, what's the real difference?
The mental model. REST says 'address a resource and act on it with a verb' (GET /users/42). RPC says 'call a function on the server' (getUser(42)). REST leans on HTTP and URLs and suits public, cacheable APIs. RPC, especially gRPC, feels like calling local code and suits fast internal service-to-service calls. Both are still request/response underneath.
Is the client ever a server too?
Yes, constantly. Client and server are roles per conversation, not permanent labels. A web server is a client to its database. A backend service is a server to the mobile app but a client to three other services. In peer-to-peer systems, every node is deliberately both.
Ways to shape the conversation
- Streaming / Server-Sent Events (SSE): the client asks once, and the server keeps sending data over time (a live score feed, a log tail). One request, many response chunks.
- WebSockets: after an initial HTTP handshake, the connection upgrades to a two-way pipe. Now the server can push to the client without a new request each time. This is how chat and live collaboration work.
| REST | RPC / gRPC | GraphQL | |
|---|---|---|---|
| Mental model | Resources at URLs + HTTP verbs | Call a function on the server | Ask for exactly the fields you want |
| Format | Usually JSON over HTTP | Often binary (Protobuf) over HTTP/2 | JSON, single flexible endpoint |
| Best at | Public APIs, caching, simplicity | Fast internal service-to-service calls | Avoiding over-/under-fetching for rich UIs |
| Watch out | Many round-trips / over-fetching | Less human-readable, tooling-heavy | Server-side complexity, harder caching |
Cutting across all of these is stateless vs stateful. A stateless server treats each request as self-contained. It remembers nothing between requests, so any copy of the server can handle any request. That's what makes horizontal scaling and load balancing easy. A stateful server keeps per-client context (an in-memory session, an open transaction), which pins a client to one server and complicates scaling and failover. The common compromise: keep servers stateless and push the state into a shared store, like a database, a cache, or a signed token the client carries.
What client-server buys, and what it costs
| Client-server | Peer-to-peer (P2P) | |
|---|---|---|
| Roles | Fixed: clients ask, server answers | Symmetric: every node is both client and server |
| Source of truth | Central, the server | Distributed across peers |
| Simplicity | Simple to build, secure, and reason about | Complex: discovery, trust, coordination |
| Bottleneck / SPOF | The server limits scale and can take everyone down | No single bottleneck; resilient to one node failing |
| Examples | The web, mobile apps + APIs, databases, email | BitTorrent, blockchains, Kademlia DHTs |
References
- MDN — Client-server overview — the gentlest intro to requests, responses, and HTTP methods
- MDN — An overview of HTTP — messages, methods, status codes, and the HTTP flow
- RFC 9110 — HTTP Semantics — the authoritative spec for what requests and responses mean
- gRPC — Introduction to gRPC — the RPC alternative to REST: call a remote method like a local one