HotShard
Caching patterns

Caching Strategies

How a cache stays fast, and where it quietly disagrees with the database behind it.

A cache is a small, fast copy of data that sits in front of a slower database. The cache itself is easy. The hard part is the policy: who loads the cache on a miss, how a write reaches the database, how long an entry lives, and what happens when a popular entry expires. Those choices decide whether the cache speeds things up safely or serves stale data and melts your database. (How a full cache decides what to evict is the LRU/LFU cache page.)

~7 min read

The problem: a fast copy that can quietly disagree#

TL;DRthe 30-second version
  • The cache is a fast copy. The backing store is the slow source of truth. A caching strategy is the rule for who reads, who writes, and in what order, so the two stay close enough.
  • Read patterns decide who loads on a miss. Cache-aside puts the app in charge. Read-through hides the load inside the cache.
  • Write patterns decide how a change reaches the store. Write-through writes both, safe but slow. Write-back writes the cache now and the store later, fast but a crash loses data. Write-around writes the store only. Write-invalidate writes the store, then deletes the cached copy.
  • The three load-spike failures are the stampede, penetration, and the avalanche. The fixes are request coalescing, negative caching, and jittered TTLs.
  • There is no best strategy. Each one trades read latency, write latency, consistency, and durability. Pick by the read/write mix and how much staleness you can accept.

Say your database serves a few thousand reads a second and your traffic wants a hundred thousand. A cache absorbs most of those reads from memory in well under a millisecond. But now the same data lives in two places. The moment they can disagree, you have a correctness problem sitting on top of a performance win.

So caching is really a set of policy choices about three actors: the application, which wants answers; the cache, a fast copy that can vanish; and the backing store, the slow source of truth. On a read miss, who loads the data and who fills the cache? On a write, do you update the cache, the store, or both, and in what order? How long does a cached entry live? And what happens when a popular key expires and a thousand requests miss at once?

appreads and writes
hit: serve · miss: go downcache-aside fills the cache on a miss
cachefast · a copy (RAM, ~ms)
load / flush / invalidate
backing storeslow · the source of truth (DB / object store, ~10–100ms)
The three actors and the data flow between them

Read patterns: cache-aside vs read-through#

Cache-aside, also called lazy loading, puts the application in charge. The app checks the cache. On a miss it reads the store itself, then writes the value into the cache so the next read hits. It's the most common pattern because it's simple and the cache and store stay decoupled. If the cache is down, the app just talks to the store directly. Slower, but not broken. The cost is that the loading logic lives in every caller.

  1. App asks the cache for key K.
  2. Hit: return the cached value. Done in about 1 ms.
  3. Miss: the app reads K from the backing store.
  4. The app writes K into the cache, usually with a TTL.
  5. The app returns the value. The next read for K is a hit.

Read-through moves that logic into the cache layer. The app only ever talks to the cache. On a miss the cache loads from the store, keeps a copy, and returns the value. The behaviour is the same as cache-aside, but the loading code lives in one place, a cache library or provider, instead of in every caller. That also lets the cache library merge concurrent misses for the same key into one load. One limit: a read-through cache usually needs the cached object to map cleanly to a store entity. Cache-aside can cache anything, a joined result or a rendered fragment.

Either way, the first read is slowBoth fill the cache only after a miss. So the first request for a key pays the full store latency, and the reads after it are fast hits until the entry expires or is evicted.
PredictWith cache-aside, two requests miss on the same cold key at the same instant. What happens?

Hint: Who loads on a miss, and is there any coordination between callers?

Both load from the store and both write the cache. That's a duplicate load. For two requests it's harmless: the values are the same and the last write wins. But if thousands of requests miss the same hot key at once, it becomes a stampede that can overwhelm the store. The fix is request coalescing, also called single-flight: one request loads while the others wait and share the result.

Write patterns: through, back, around, invalidate#

Read patterns decide who fills the cache. Write patterns decide how a change reaches the store, and what happens to the cached copy. There are four common choices. They differ in where the latency lands and what a crash can lose.

  • Write-through: write the cache and the store together, and acknowledge only when both are done. The cache always agrees with the store and reads are warm. But every write pays the store's latency. Best when reads follow writes soon and you can't serve a stale value.
  • Write-back (also called write-behind): write only the cache, mark the entry dirty, and acknowledge at once. Dirty entries flush to the store later, in batches. Writes are very fast and the store sees fewer, merged writes. But a crash before the flush silently loses every write that hasn't flushed yet.
  • Write-around: write straight to the store and skip the cache. The cache fills only on a later read miss. Good for data that is written a lot and rarely read soon, like logs and audit records. It keeps entries nobody will read out of the cache, but the first read after a write is always a miss.
  • Write-invalidate: write the store, then delete the cached entry instead of updating it. The next read reloads a fresh copy. This is the usual write path paired with cache-aside reads.
Why delete instead of update?Updating the cache on a write looks natural, but it invites a race. Two writers can interleave their cache update and their store update, so the cache ends up holding the older value. A delete can't go wrong that way. Run it twice and nothing changes, and the next reader simply reloads the current value from the store. Invalidate is the safer default.

These compose. Write-through with read-through gives an always-warm cache that agrees with the store. Cache-aside with write-invalidate gives a lazy cache that is stale for at most the TTL.

The numbers: latency, hit rate, staleness window#

One equation governs the payoff. Call the cache hit latency h, the store latency m, and the hit rate r. The average read latency is r·h + (1−r)·(m + h). A miss pays the store and then the cache write, which is the (m + h). Because m is often 10 to 100 times h, the hit rate dominates everything.

  • Hit rate is non-linear in value. Going from 90% to 99% cuts the share of requests that reach the store by about ten times. That is often the difference between the database coping and falling over.
  • A RAM cache hit is about 0.1 to 1 ms. A database miss is about 5 to 50 ms. A cross-region or disk miss can be 100 ms or more.
  • Store load is what you're really protecting. If the store can handle Q queries a second and you receive N reads a second, you need a hit rate of at least 1 − Q/N or the store saturates. The cache bounds store load as much as it cuts latency.
  • Hit rate is capped by how much of the hot set fits in memory. If 20% of keys serve 80% of reads, which is common, caching that 20% gets you most of the benefit. Pushing higher means caching the long tail, which costs a lot more memory for less and less return.
The staleness windowWith TTL expiry, the worst-case staleness is the TTL. A value changed in the store right after a cache fill can be served stale for up to TTL seconds. With write-invalidate, the window shrinks to how long the delete takes to reach the cache. Choosing the TTL is choosing how stale you're willing to be.
The core tension: consistency vs latency vs durability

Every strategy is a point in a triangle: consistency (does the cache agree with the store?), latency (how fast are reads and writes?), and durability (can an acknowledged write be lost?). You can't have all three at once.

  • Consistency vs latency: write-through and write-invalidate keep the cache close to the store, but add a synchronous store hop to every write. Cache-aside with a long TTL is fast, but can serve data that is stale by up to the TTL.
  • Latency vs durability: write-back gives the fastest writes because it acknowledges before the store is updated. That unflushed window is exactly what a crash loses. Write-through is durable because the store is written before the ack.
  • Consistency across servers: when many app servers each hold their own cache, a write must invalidate all of them. That needs an invalidation message bus, or short TTLs as a crude substitute. There is always a propagation delay during which different servers can disagree.
Pick by workloadRead-heavy and staleness-tolerant (feeds, catalogs): cache-aside plus TTL. Write-then-read and must be fresh (profile edits): write-through. Write-heavy and rarely read (logs, events): write-around. Throughput-critical and loss-tolerant (counters, metrics buffers): write-back with a persistent log to bound the loss.
Strategies side by side
StrategyConsistencyRead/write latencyDurabilityBest for
Cache-aside (lazy)Eventual, bounded by TTLFast reads on hit; a miss pays the store; fast writesStore is durable; losing the cache is fineRead-heavy, staleness-tolerant; the default
Read-throughEventual, bounded by TTLSame as cache-aside; load code in one placeStore is durableRead-heavy, with a cache library that loads for you
Write-throughStrong (cache = store)Slow writes (cache + store together); warm readsDurable; store written before the ackWrite-then-read, must-be-fresh data
Write-back / behindWeak until the flushFastest writes; warm readsAt risk; a crash loses unflushed writesThroughput-critical, loss-tolerant (counters)
Write-aroundFresh in the store; cache fills on readFast writes; first read is a missDurable; store written directlyWrite-heavy, rarely read soon (logs, events)
Where these patterns run in the wild
  • Redis and Memcached are the two default in-memory caches. Cache-aside plus write-invalidate on Redis is the most common application caching pattern.
  • Facebook memcache is the canonical large-scale deployment ('Scaling Memcache at Facebook', NSDI 2013). It is cache-aside, which they call demand-filled look-aside, with a region-wide invalidation pipeline driven off the database commit log. A miss hands one client a lease, a short-lived token, and only that client loads from the store and sets the value. Other clients that miss the same key wait briefly and retry, so the store sees one load. If the key was invalidated while that client was loading, its lease is voided and its stale value is rejected. One mechanism kills both the stampede and the stale set.
  • Netflix EVCache is a replicated caching tier built on Memcached, spanning AWS availability zones. Writes fan out to several zones so it survives a zone failure at very high hit rates.
  • CDNs (CloudFront, Fastly, Cloudflare) cache at the network edge. They live by TTLs in Cache-Control headers, stale-while-revalidate (keep serving the expired value while one background load refreshes it), and request coalescing at each edge so the origin sees one fetch per object, not one per viewer.
Failure modes: stampede, penetration, avalanche, loss

Caches fail in characteristic ways, almost always by letting too much traffic reach the store at once. Each failure has a name and a specific fix.

  • Cache stampede (thundering herd): one hot key expires or is evicted, and every concurrent request for it misses at the same instant. They all hit the store with the same query. Fix: request coalescing (one loader, the rest wait and share the result), a short lock or lease per key, or refreshing a hot key slightly before it expires.
  • Cache penetration: requests for keys that don't exist anywhere skip the cache, because there is nothing to hit, and reach the store every time. Common with buggy or malicious clients probing random IDs. Fix: negative caching (cache the 'not found' for a short TTL) or a Bloom filter in front of the cache that rejects keys that definitely don't exist.
  • Cache avalanche: a large number of entries expire at the same moment, for example everything loaded at startup with the same one-hour TTL, so a wave of misses hits the store at once. Fix: jitter the TTLs (add a small random spread, say 3600 seconds plus or minus a few hundred) and warm the cache gradually.
  • Write-back data loss: every write that is acknowledged but not yet flushed lives only in cache memory. A crash, restart, or eviction of that dirty entry before the flush loses it silently. The exposure is the flush interval times the write rate. Mitigate with a persistent log for the dirty buffer, replication of the cache tier, or by using write-back only for data you can afford to lose.
  • Stale reads: any pattern that fills the cache and then lets the store change underneath it (a long TTL, a failed invalidation, replication lag) serves data older than the source of truth. Bound it with shorter TTLs, reliable invalidation, or versioned keys.
In an interview

Frame it as three actors and two sets of choices. Name the read patterns and the write patterns, pick one for the workload, and say why you invalidate rather than update the cache. Then volunteer the failure modes and their fixes. That is what separates a strong answer.

Cache-aside vs read-through: what's the actual difference?

Both load on a miss. The difference is where the load logic lives. Cache-aside puts it in the application: the app checks the cache, and on a miss it reads the store and fills the cache itself. Read-through puts it in the cache layer: the app only talks to the cache, and the cache loads from the store on a miss. Cache-aside is more flexible and lets the app fall back to the store if the cache is down. Read-through keeps the logic in one place and can merge concurrent misses for you.

When does write-back actually lose data?

Write-back acknowledges a write as soon as it's in the cache and flushes to the store later. Any write that has been acknowledged but not yet flushed lives only in cache memory. A crash, restart, or eviction before the flush loses it. Mitigate with a persistent log for the dirty buffer, replicate the cache tier, or use write-back only for data you can afford to lose.

What's a cache stampede and how do you stop it?

A hot key expires and every concurrent request for it misses at the same instant, so they all hit the store with the same query. The main fix is request coalescing, also called single-flight: exactly one request loads the value while the others wait and share the result. Back it up with a per-key lock or lease (Facebook's approach), TTL jitter so keys don't expire together, and refreshing a hot key shortly before it expires.

How do you choose the TTL?

A shorter TTL means fresher data but a lower hit rate and more store load, because entries expire and reload more often. A longer TTL means a higher hit rate but you can serve data up to TTL seconds out of date. Whatever you choose, jitter it. Identical TTLs on many keys cause an avalanche when they all expire at once.

Should I update the cache on a write, or just delete it?

Prefer deleting. Updating on write opens a race where two writers interleave and the cache ends up holding the older value. Deleting can't go wrong that way, and the next read simply reloads the current value from the store. Update-in-place is only worth it when reloads are expensive and you've handled the ordering carefully.

References & further reading
References

Feedback on this topic →