Start here: memory is fast, but a crash erases it#
TL;DRthe 30-second version
- A database promises durability, the D in ACID: an acknowledged write survives a crash. But its fast working copy lives in memory, and a crash wipes memory.
- Writing each change straight into the on-disk data file is durable but slow. The writes land at scattered places, and a crash halfway through a page can corrupt it.
- The write-ahead log fixes this. Append the change to a log file on disk first, make it durable, then apply it to memory and acknowledge. Log before apply.
- After a crash, replay the log to redo every logged change. A periodic checkpoint writes memory to the data file and truncates the log, so recovery stays short.
The trouble is where the database does its fast work. It keeps the live copy of the data in memory (RAM), and memory only holds data while the machine is powered. Pull the plug and everything in it is gone. So a write that exists only in memory isn't durable. Disk is the opposite: an SSD or hard drive keeps its contents across a crash.
The obvious fix is to write each change straight to its home in the on-disk data file before acknowledging. That's safe but slow, for two reasons. Records sit at scattered positions in the file, so the disk jumps around. That's random I/O, the slowest thing a disk does. Worse, updating data in place isn't crash-safe. Disks write in fixed-size blocks called pages. If the machine dies halfway through overwriting a page, the page is left half old and half new. That's a torn write, and now the data file itself is corrupt.
The fix: write to a log before you touch anything#
The write-ahead log is an append-only file on disk. You only ever add to the end; you never go back and overwrite. When a change comes in, the database does three things in a strict order.
- Append a record describing the change to the end of the log, for example "set user:1 to alice". Each record carries its length, so recovery can read them back one at a time.
- Call fsync. That's an operating-system request that tells the disk to actually persist the bytes, not just buffer them. It doesn't return until they're on the drive. From here on the record survives a crash.
- Only then apply the change to the in-memory table that queries read from, and tell the client "committed."
Why is this faster than writing the data file directly? Because appending to a log is sequential I/O. Every write lands at the end of one file, so the disk never seeks to scattered locations. Sequential writes are far faster than the random writes an in-place update needs. You still pay one durable write per change, but you pay it in the cheapest form.
PredictThe database appends a write to the log, fsync returns, and it acknowledges the client. Then it crashes before applying the change to its in-memory table. What happens to that write after recovery?
Hint: What was on disk at the moment the client heard "committed"?
It's recovered. The log record was durable before the client was acknowledged, so it's on disk even though memory never got the change. On restart the database reads the log and redoes the change. This is replay. The in-memory state is disposable; the log is the source of truth for anything not yet in the data file.
Checkpoints: keeping the log from growing forever#
If every change appends a record and nothing is ever removed, the log grows without bound, and replay takes longer and longer. A checkpoint fixes that.
- Write the current in-memory state to the on-disk data file, its permanent home.
- Record in the log that a checkpoint happened here.
- Truncate the log. Every record before the checkpoint is now in the data file, so recovery would never need it.
- Recovery after a crash is now: reload the data file, then replay only the records since the last checkpoint.
What it costs: writing twice#
A WAL isn't free. Every change hits disk twice: once as a log record, and again when a checkpoint writes it into the data file. That's write amplification, roughly 2× the logical data volume. The second write is deferred, so many changes to the same page collapse into one physical write at checkpoint time.
The other cost is the fsync on the write path. One fsync per write would cap throughput hard. The standard escape is group commit: gather the log records of many concurrent transactions and flush them with a single fsync. One durable write, many transactions made safe.
| Approach | Speed | Crash safety |
|---|---|---|
| Write to memory only | Fastest | None — a crash loses everything not in the data file |
| Update the data file in place per write | Slow (random I/O) | Poor — a torn write can corrupt a page |
| Write-ahead log + checkpoint | Fast (sequential log, batched fsync) | Full — replay restores every acknowledged write |
The durability knob#
"Log before apply" fixes the order. It doesn't say when the client hears "committed" relative to the fsync. That timing is a knob, and it trades latency against how much a crash can lose.
| Setting | When it acknowledges | Crash loses | Cost |
|---|---|---|---|
| Synchronous commit | After the log fsync completes | Nothing acknowledged | Higher write latency |
| Group commit | After a shared batch fsync | Nothing acknowledged | Tiny added latency, much higher throughput |
| Asynchronous commit | Before the fsync, trusting it will land | The last few milliseconds of "committed" writes | Lowest latency, a small loss window |
Asynchronous commit is much faster, but a crash in that window loses writes the client thinks are safe. That's a fine trade for data you can regenerate and a terrible one for money. Redis exposes exactly this dial for its append-only file: fsync always, every second, or when the OS decides.
If this comes up in an interview#
What if the crash happens while the log record itself is half-written?
That partial record is at the very end of the log, and no client was ever told "committed" for it, because fsync never returned. Recovery sees the torn tail (records carry a length, often a checksum too) and discards it. You lose only a write that was never promised.
Where does this show up in real systems?
PostgreSQL's WAL, which also feeds its replicas. MySQL InnoDB's redo log. Redis's append-only file, replayed on startup to rebuild memory. SQLite's WAL mode. And an LSM tree's memtable is guarded by a WAL, since its buffered writes would otherwise vanish on a crash.
How is this different from a journaling filesystem?
Same idea at a different layer. A journaling filesystem write-ahead-logs its own metadata so the filesystem structure survives a crash. A database can't rely on that for its own durability, so it keeps its own log, sometimes on top of a journaling filesystem.
References
- PostgreSQL — Write-Ahead Logging (WAL) — The canonical, readable explanation of WAL, checkpoints, and commit settings.
- SQLite — Write-Ahead Logging — How SQLite's WAL mode works and why it improves concurrency.
- Redis — Persistence (AOF) — The append-only file and its fsync policies — the durability knob in practice.
- ARIES (Mohan et al., 1992) — The foundational paper on write-ahead logging, checkpoints, and crash recovery.