The problem: a fast copy that can quietly disagree#
TL;DRthe 30-second version
- The cache is a fast copy. The backing store is the slow source of truth. A caching strategy is the rule for who reads, who writes, and in what order, so the two stay close enough.
- Read patterns decide who loads on a miss. Cache-aside puts the app in charge. Read-through hides the load inside the cache.
- Write patterns decide how a change reaches the store. Write-through writes both, safe but slow. Write-back writes the cache now and the store later, fast but a crash loses data. Write-around writes the store only. Write-invalidate writes the store, then deletes the cached copy.
- The three load-spike failures are the stampede, penetration, and the avalanche. The fixes are request coalescing, negative caching, and jittered TTLs.
- There is no best strategy. Each one trades read latency, write latency, consistency, and durability. Pick by the read/write mix and how much staleness you can accept.
Say your database serves a few thousand reads a second and your traffic wants a hundred thousand. A cache absorbs most of those reads from memory in well under a millisecond. But now the same data lives in two places. The moment they can disagree, you have a correctness problem sitting on top of a performance win.
So caching is really a set of policy choices about three actors: the application, which wants answers; the cache, a fast copy that can vanish; and the backing store, the slow source of truth. On a read miss, who loads the data and who fills the cache? On a write, do you update the cache, the store, or both, and in what order? How long does a cached entry live? And what happens when a popular key expires and a thousand requests miss at once?
Read patterns: cache-aside vs read-through#
Cache-aside, also called lazy loading, puts the application in charge. The app checks the cache. On a miss it reads the store itself, then writes the value into the cache so the next read hits. It's the most common pattern because it's simple and the cache and store stay decoupled. If the cache is down, the app just talks to the store directly. Slower, but not broken. The cost is that the loading logic lives in every caller.
- App asks the cache for key K.
- Hit: return the cached value. Done in about 1 ms.
- Miss: the app reads K from the backing store.
- The app writes K into the cache, usually with a TTL.
- The app returns the value. The next read for K is a hit.
Read-through moves that logic into the cache layer. The app only ever talks to the cache. On a miss the cache loads from the store, keeps a copy, and returns the value. The behaviour is the same as cache-aside, but the loading code lives in one place, a cache library or provider, instead of in every caller. That also lets the cache library merge concurrent misses for the same key into one load. One limit: a read-through cache usually needs the cached object to map cleanly to a store entity. Cache-aside can cache anything, a joined result or a rendered fragment.
PredictWith cache-aside, two requests miss on the same cold key at the same instant. What happens?
Hint: Who loads on a miss, and is there any coordination between callers?
Both load from the store and both write the cache. That's a duplicate load. For two requests it's harmless: the values are the same and the last write wins. But if thousands of requests miss the same hot key at once, it becomes a stampede that can overwhelm the store. The fix is request coalescing, also called single-flight: one request loads while the others wait and share the result.
Write patterns: through, back, around, invalidate#
Read patterns decide who fills the cache. Write patterns decide how a change reaches the store, and what happens to the cached copy. There are four common choices. They differ in where the latency lands and what a crash can lose.
- Write-through: write the cache and the store together, and acknowledge only when both are done. The cache always agrees with the store and reads are warm. But every write pays the store's latency. Best when reads follow writes soon and you can't serve a stale value.
- Write-back (also called write-behind): write only the cache, mark the entry dirty, and acknowledge at once. Dirty entries flush to the store later, in batches. Writes are very fast and the store sees fewer, merged writes. But a crash before the flush silently loses every write that hasn't flushed yet.
- Write-around: write straight to the store and skip the cache. The cache fills only on a later read miss. Good for data that is written a lot and rarely read soon, like logs and audit records. It keeps entries nobody will read out of the cache, but the first read after a write is always a miss.
- Write-invalidate: write the store, then delete the cached entry instead of updating it. The next read reloads a fresh copy. This is the usual write path paired with cache-aside reads.
These compose. Write-through with read-through gives an always-warm cache that agrees with the store. Cache-aside with write-invalidate gives a lazy cache that is stale for at most the TTL.
The numbers: latency, hit rate, staleness window#
One equation governs the payoff. Call the cache hit latency h, the store latency m, and the hit rate r. The average read latency is r·h + (1−r)·(m + h). A miss pays the store and then the cache write, which is the (m + h). Because m is often 10 to 100 times h, the hit rate dominates everything.
- Hit rate is non-linear in value. Going from 90% to 99% cuts the share of requests that reach the store by about ten times. That is often the difference between the database coping and falling over.
- A RAM cache hit is about 0.1 to 1 ms. A database miss is about 5 to 50 ms. A cross-region or disk miss can be 100 ms or more.
- Store load is what you're really protecting. If the store can handle Q queries a second and you receive N reads a second, you need a hit rate of at least 1 − Q/N or the store saturates. The cache bounds store load as much as it cuts latency.
- Hit rate is capped by how much of the hot set fits in memory. If 20% of keys serve 80% of reads, which is common, caching that 20% gets you most of the benefit. Pushing higher means caching the long tail, which costs a lot more memory for less and less return.
The core tension: consistency vs latency vs durability
Every strategy is a point in a triangle: consistency (does the cache agree with the store?), latency (how fast are reads and writes?), and durability (can an acknowledged write be lost?). You can't have all three at once.
- Consistency vs latency: write-through and write-invalidate keep the cache close to the store, but add a synchronous store hop to every write. Cache-aside with a long TTL is fast, but can serve data that is stale by up to the TTL.
- Latency vs durability: write-back gives the fastest writes because it acknowledges before the store is updated. That unflushed window is exactly what a crash loses. Write-through is durable because the store is written before the ack.
- Consistency across servers: when many app servers each hold their own cache, a write must invalidate all of them. That needs an invalidation message bus, or short TTLs as a crude substitute. There is always a propagation delay during which different servers can disagree.
Strategies side by side
| Strategy | Consistency | Read/write latency | Durability | Best for |
|---|---|---|---|---|
| Cache-aside (lazy) | Eventual, bounded by TTL | Fast reads on hit; a miss pays the store; fast writes | Store is durable; losing the cache is fine | Read-heavy, staleness-tolerant; the default |
| Read-through | Eventual, bounded by TTL | Same as cache-aside; load code in one place | Store is durable | Read-heavy, with a cache library that loads for you |
| Write-through | Strong (cache = store) | Slow writes (cache + store together); warm reads | Durable; store written before the ack | Write-then-read, must-be-fresh data |
| Write-back / behind | Weak until the flush | Fastest writes; warm reads | At risk; a crash loses unflushed writes | Throughput-critical, loss-tolerant (counters) |
| Write-around | Fresh in the store; cache fills on read | Fast writes; first read is a miss | Durable; store written directly | Write-heavy, rarely read soon (logs, events) |
Where these patterns run in the wild
- Redis and Memcached are the two default in-memory caches. Cache-aside plus write-invalidate on Redis is the most common application caching pattern.
- Facebook memcache is the canonical large-scale deployment ('Scaling Memcache at Facebook', NSDI 2013). It is cache-aside, which they call demand-filled look-aside, with a region-wide invalidation pipeline driven off the database commit log. A miss hands one client a lease, a short-lived token, and only that client loads from the store and sets the value. Other clients that miss the same key wait briefly and retry, so the store sees one load. If the key was invalidated while that client was loading, its lease is voided and its stale value is rejected. One mechanism kills both the stampede and the stale set.
- Netflix EVCache is a replicated caching tier built on Memcached, spanning AWS availability zones. Writes fan out to several zones so it survives a zone failure at very high hit rates.
- CDNs (CloudFront, Fastly, Cloudflare) cache at the network edge. They live by TTLs in Cache-Control headers, stale-while-revalidate (keep serving the expired value while one background load refreshes it), and request coalescing at each edge so the origin sees one fetch per object, not one per viewer.
Failure modes: stampede, penetration, avalanche, loss
Caches fail in characteristic ways, almost always by letting too much traffic reach the store at once. Each failure has a name and a specific fix.
- Cache stampede (thundering herd): one hot key expires or is evicted, and every concurrent request for it misses at the same instant. They all hit the store with the same query. Fix: request coalescing (one loader, the rest wait and share the result), a short lock or lease per key, or refreshing a hot key slightly before it expires.
- Cache penetration: requests for keys that don't exist anywhere skip the cache, because there is nothing to hit, and reach the store every time. Common with buggy or malicious clients probing random IDs. Fix: negative caching (cache the 'not found' for a short TTL) or a Bloom filter in front of the cache that rejects keys that definitely don't exist.
- Cache avalanche: a large number of entries expire at the same moment, for example everything loaded at startup with the same one-hour TTL, so a wave of misses hits the store at once. Fix: jitter the TTLs (add a small random spread, say 3600 seconds plus or minus a few hundred) and warm the cache gradually.
- Write-back data loss: every write that is acknowledged but not yet flushed lives only in cache memory. A crash, restart, or eviction of that dirty entry before the flush loses it silently. The exposure is the flush interval times the write rate. Mitigate with a persistent log for the dirty buffer, replication of the cache tier, or by using write-back only for data you can afford to lose.
- Stale reads: any pattern that fills the cache and then lets the store change underneath it (a long TTL, a failed invalidation, replication lag) serves data older than the source of truth. Bound it with shorter TTLs, reliable invalidation, or versioned keys.
In an interview
Frame it as three actors and two sets of choices. Name the read patterns and the write patterns, pick one for the workload, and say why you invalidate rather than update the cache. Then volunteer the failure modes and their fixes. That is what separates a strong answer.
Cache-aside vs read-through: what's the actual difference?
Both load on a miss. The difference is where the load logic lives. Cache-aside puts it in the application: the app checks the cache, and on a miss it reads the store and fills the cache itself. Read-through puts it in the cache layer: the app only talks to the cache, and the cache loads from the store on a miss. Cache-aside is more flexible and lets the app fall back to the store if the cache is down. Read-through keeps the logic in one place and can merge concurrent misses for you.
When does write-back actually lose data?
Write-back acknowledges a write as soon as it's in the cache and flushes to the store later. Any write that has been acknowledged but not yet flushed lives only in cache memory. A crash, restart, or eviction before the flush loses it. Mitigate with a persistent log for the dirty buffer, replicate the cache tier, or use write-back only for data you can afford to lose.
What's a cache stampede and how do you stop it?
A hot key expires and every concurrent request for it misses at the same instant, so they all hit the store with the same query. The main fix is request coalescing, also called single-flight: exactly one request loads the value while the others wait and share the result. Back it up with a per-key lock or lease (Facebook's approach), TTL jitter so keys don't expire together, and refreshing a hot key shortly before it expires.
How do you choose the TTL?
A shorter TTL means fresher data but a lower hit rate and more store load, because entries expire and reload more often. A longer TTL means a higher hit rate but you can serve data up to TTL seconds out of date. Whatever you choose, jitter it. Identical TTLs on many keys cause an avalanche when they all expire at once.
Should I update the cache on a write, or just delete it?
Prefer deleting. Updating on write opens a race where two writers interleave and the cache ends up holding the older value. Deleting can't go wrong that way, and the next read simply reloads the current value from the store. Update-in-place is only worth it when reloads are expensive and you've handled the ordering carefully.
References & further reading
- Nishtala et al. — Scaling Memcache at Facebook (NSDI 2013) — leases, stale-set protection, and region-wide invalidation at scale
- AWS — Database Caching Strategies Using Redis (whitepaper) — cache-aside, write-through, TTL, and when to use each
- AWS ElastiCache — Caching strategies (lazy loading, write-through, TTL) — pseudocode and trade-offs for the core read/write patterns
- Netflix Tech Blog — Announcing EVCache — a replicated, multi-AZ Memcached caching tier