HotShard
Operating systems

The Page Cache

How the operating system keeps file pages in RAM, defers writes, and pays for it on eviction.

Every file read or write goes through the kernel's page cache: a pool of RAM holding recently-touched pieces of files. A read that finds its page in RAM never touches the disk. A write lands in RAM and returns at once, with the disk update deferred. That is why the second open of a large file is instant. This page is about the one thing a plain cache doesn't have: the difference between a clean page and a dirty one.

~5 min read

Start here: disk is slow and RAM is a bounded copy#

TL;DRthe 30-second version
  • Disk is slow. The page cache keeps recently-used pieces of files (pages, usually 4 KB) in RAM so most reads and writes hit memory.
  • A read that finds its page in RAM is a hit. A miss faults the page in from disk into a free frame, a slot of RAM that holds one page.
  • A write updates the page in RAM, marks it dirty, and returns. A writeback later flushes dirty pages to disk and marks them clean.
  • When RAM is full, a clean page can be dropped for free, because disk already has the same bytes. A dirty page must be written back first, or the write is lost.

Reading a byte from RAM takes about 100 nanoseconds. Reading it from a solid-state disk takes tens of microseconds, and from a spinning disk, milliseconds. That is a gap of a hundred to ten thousand times. If every file access paid the disk price, no program would feel fast. So the kernel keeps copies of file pages it has touched recently in spare RAM, betting they will be touched again soon.

programread() / write() a file
hit: serve from RAM · miss: fault in ↓
page cachepages in RAM frames · clean or dirty · ~100 ns
writeback ↑ flushes dirty pages
diskthe durable file · ~10 µs–10 ms
Where the page cache sits

How it works: fault in, mark dirty, write back, evict#

A read first checks whether the page is already in a frame. If it is, that is a hit: the kernel copies the bytes out of RAM and no disk access happens. If not, that is a miss. The kernel finds a free frame, reads the 4 KB page from disk into it, and then serves your bytes. The page is now clean, meaning its RAM copy matches disk exactly.

A write updates the bytes in RAM and marks the frame dirty: RAM now holds newer data than disk. The write returns immediately, without waiting for the disk. If the page wasn't resident, the kernel faults it in first, then applies the write. Either way, disk is stale for that page until a writeback catches it up.

A writeback walks the dirty pages, writes each one to disk, and marks it clean again. The page stays in the cache. In a real kernel this runs in the background on a timer, and on demand when you call fsync. The payoff is batching: if you wrote the same page a thousand times, one writeback settles all thousand changes with one disk write.

Now the part a plain cache never has to think about: eviction when the cache is full. The kernel picks a victim, roughly the least-recently-used page. If the victim is clean, reclaiming it is free. Disk already holds the same bytes, so the kernel just drops the frame. If the victim is dirty, it is the only up-to-date copy of that data anywhere. The kernel must write it back to disk first, then free the frame. That forced disk write is the hidden cost of a dirty page, and it is why the kernel prefers clean victims when it can.

on disknot cached
read miss
cleanfaulted in by a read
write
dirtya write landed
writeback / fsync
clean againwriteback flushed it
The life of a page
PredictYou write to a page, then immediately read it back before any writeback runs. Does the read hit or miss, and does it return the old disk value or your new one?

Hint: A read only touches disk on a miss. The written page is still resident.

It is a hit, and it returns your new value. The write updated the page in RAM and marked it dirty. The read finds that same resident frame and serves the RAM copy, which is the newest. Disk still holds the old value, but nobody reads disk on a hit.

If this comes up in an interview#

The one-linerThe page cache is an LRU cache of file pages in RAM with one extra bit: a dirty page owes disk a write, so evicting it costs a writeback, while a clean page is dropped for free.
What is the difference between a clean and a dirty page?

A clean page's RAM copy matches disk exactly, so it can be evicted for free by dropping the frame. A dirty page holds writes disk has not seen yet, so it is the only current copy. It must be written back before its frame is reused, or the data is lost.

Why do writes return before the data is on disk?

The write only updates the page in RAM and marks it dirty. The disk update is deferred to a later writeback. That makes writes fast and lets many writes to the same page batch into one disk write. The cost is a durability window: an un-flushed write is lost on a crash unless you fsync.

The cache is full and you need to bring in a new page. What happens?

The kernel evicts roughly the least-recently-used page, preferring a clean victim. A clean page is dropped for free. If it must evict a dirty page, it writes that page to disk first, then frees the frame. That is an extra disk write on the eviction path.

How would you guarantee a write survives a crash?

Call fsync, which forces a writeback of the file's dirty pages and waits for the disk to confirm. Until that returns, the write is only in volatile RAM. Databases build their durability on exactly this, paired with a write-ahead log.

The trade-offs
  • Speed versus durability: writes return the instant they hit RAM, but they are not safe until a writeback reaches disk. The kernel bounds the window with a timer (dirty pages older than a few seconds get flushed) and a ceiling on how much of RAM may be dirty. fsync trades the speed back for a guarantee.
  • Batching versus latency: letting dirty pages pile up lets many writes settle in one flush, but a bigger pile means a longer, spikier flush and more data at risk. On Linux, vm.dirty_background_ratio starts background flushing at a low-water mark and vm.dirty_ratio blocks writers at a high one.
  • Cache size versus application memory: the cache uses free RAM and shrinks when programs need memory. A memory-hungry program can push hot file pages out and slow every later read.
  • Bypassing it: O_DIRECT skips the page cache entirely, so a database can manage its own buffer pool without keeping two copies of the same page.
Common pitfalls
  • Assuming a completed write() is durable. It is only in RAM until a writeback. If durability matters, call fsync and check that it succeeded.
  • Reading high cache usage as a memory leak. On Linux, the 'buff/cache' column in free -m is the page cache. Clean pages are reclaimable, and the kernel drops them instantly under pressure. Low free memory with a large cache is normal and healthy.
  • Benchmarking with a warm cache. The 'first run slow, second run fast' effect on builds and test suites is the cache warming up. If you do not drop caches (on Linux, /proc/sys/vm/drop_caches), your second run measures RAM, not disk.
  • Ignoring the dirty ceiling. A write-heavy job that outruns the disk fills the dirty budget and then throttles to disk speed, often as a sudden stall.
  • Double caching by accident. PostgreSQL deliberately keeps a modest buffer pool and leans on the OS page cache. A database that manages its own large buffer pool without O_DIRECT (some InnoDB configurations) can hold two copies of the same page and waste RAM.
References
References

Feedback on this topic →