Databases / Working draft
Keeping a copy
of the answer
A cache saves repeated work. Keeping its answers useful takes a little more thought.
You run a shop selling ceramics. The blue-mug page shows a photograph, a price of €18, and four mugs in stock. To assemble that page, the application fetches the product and its current price and inventory from the catalogue database. The next visitor asks for exactly the same information. And so does the next.
The database may already keep its disk pages in memory, and an index may make this lookup cheap. Even so, the application repeats a query, transfers its result, and assembles an answer. If that repeated work matters, we could keep the finished product response somewhere the application can retrieve it directly.
That saved answer is a cache entry. Its key might be product:blue-mug:EUR; its value contains the name, price, and stock count. We've introduced another copy of information. Now we have two questions: how much work does the copy save, and when can we trust it?
The first visitor still does the work
With cache-aside, the application looks for the key first. An absent key is a miss. On a miss, it queries the catalogue, saves the result in the cache, and sends the result to the visitor. A later request that finds the key is a hit: the application sends the saved value without repeating the catalogue query.
Here, the cache is a separate service, shared by the application's servers. A local in-process cache could remove the network lookup too, but each server would have its own copies to maintain. Redis and Valkey are examples of services that can hold keyed values in memory. Using one doesn't automatically provide cache-aside: the application still needs to decide what to save and when to fetch it.
The key must describe the answer completely. A euro price and a dollar price aren't interchangeable, nor are public prices and a customer's negotiated discount. Omitting currency or customer scope can make a very fast cache return the wrong answer. Caching a product response also creates dependencies: changing either its price or its inventory may require removing that one entry.
The price changes. The copy doesn't.
The shop marks the mug down to €15. Updating the catalogue leaves the cache's €18 untouched. The cache can't infer that its value came from a particular database row, or that the row has changed. Without another step, the next hit returns the old price.
A common write path updates the catalogue first, then deletes the corresponding cache entry. This deletion is invalidation. The next reader misses and fills the cache with €15. Deleting before the catalogue update would leave an opportunity for a reader to refill it with €18 while the write is still in progress.
Even the better order isn't an atomic transaction across the two services. The application could fail after committing €15 but before deleting €18. Or a slow reader could fetch €18 before the update, then save it after the writer invalidates the key. A deletion can succeed and still be followed by an old fill. Version checks, coordinated change delivery, or a more carefully designed fill protocol may be needed when that race matters.
We can give every copy a time to live, or TTL. If the mug response was inserted at time zero with a five-minute TTL, it becomes unusable at minute five. Hits needn't extend that deadline. Shortening the TTL reduces the time an untouched copy can linger, but also causes more misses. Expiry removes an answer; it doesn't fetch a new one.
A TTL is a limit on the lifetime of this copy, rather than a universal freshness guarantee. A delayed fill can start a fresh TTL on already old data, and a query against a lagging database replica can fetch an old value again. For the product page we might accept a brief delay. At checkout, the authoritative transaction must verify the payable price and reserve stock; the displayed count can't promise that the last mug is still available.
How long does the old price stay?
Three products live in a catalogue. The cache holds two responses, initially none. Each read uses a live copy or fetches the current catalogue value. One time step represents an invented unit of time.
Read the mug twice. Turn off Invalidate on write, change its price and stock, then read it again. Advance time until the copy expires and compare the next read. Reading all three products also forces an eviction.
Controls
Result Catalogue, copies, and the last response
What this model leaves out
Requests and writes finish sequentially. The catalogue always succeeds. Copies have fixed expiry from insertion, reads do not renew it, and eviction uses exact least recently used order. Redis and Valkey offer several policies and approximate LRU. This model does not simulate concurrent fills, network delays, persistence, replication, checkout transactions, or measured speed.
There isn't room for every answer
Suppose the mug is popular but most of the catalogue is rarely visited. Keeping the mug response is useful; keeping every obscure plate response indefinitely is less useful. A cache with a memory budget needs to remove entries when it fills. That's eviction, a different reason for disappearance from expiry or invalidation.
A least-recently-used policy favours recently requested entries. A least-frequently-used policy favours entries requested often. Neither knows the future, and a new wave of traffic can upset the pattern. Redis offers configurable policies and uses sampling to approximate LRU rather than maintaining exact global recency order. A no-eviction policy instead rejects writes that need more space under its memory limit. Check which behaviour the application can tolerate.
Evicting a mug copy doesn't remove the mug from the catalogue. Its next visitor simply misses. This is why a cache should be able to lose its contents without losing the shop's records. A system holding the only copy of an order is doing a different job, even if it happens to use the same software.
An empty cache can be a busy database
If the cache is unavailable, the application can fetch the catalogue directly. That sounds reassuring until we remember how much traffic the cache was absorbing. A database handling occasional misses may struggle when every page request suddenly reaches it.
An empty restart creates a similar surge. Many requests for the mug can miss together and all begin the same query before any has filled the entry. This is often called a cache stampede. Coordinating concurrent fills for a key can let the other callers share one result. Staggered expiries, bounded fallback concurrency, and warming selected entries can also help. Each addresses a particular source of load; none increases the catalogue's capacity without limit.
Persisting cache contents can reduce rebuilding. Redis and Valkey support snapshots and append-only logs. A snapshot saves a point-in-time dataset; later changes can be absent after recovery. An append-only log records changes, with disk synchronisation policy affecting what can be lost on a crash. These mechanisms protect the cache's own state. They don't establish that a recovered €18 agrees with today's catalogue.
Replication answers another question: is there another instance with a copy? Valkey replication is asynchronous by default. Its WAIT command can wait for replica acknowledgements, but the documentation explicitly says this doesn't turn the deployment into a strongly consistent system, and acknowledged writes can still be lost in some failovers. Persistence, replication, and invalidating a product response are separate promises. Adding one doesn't silently supply the others.
What would you cache in this shop?
The public mug page is a reasonable candidate when many visitors ask for the same response and a short display delay is acceptable. An inventory reservation is different: it changes shared state and must enforce a rule. I'd keep that decision in the transactional database and treat the cached stock count as a preview.
Before adding a service, measure where the page's work goes. If the catalogue lookup is already cheap, a network cache, invalidation protocol, and failure path may cost more than they save. Track hits and misses alongside latency, errors, evictions, stale responses, and the load of rebuilding. A high hit rate alone won't tell us whether checkout is correct.
Now imagine a price varies for every signed-in customer. The shared product key no longer identifies the answer. Separate customer keys could be correct, but there may be too little reuse to justify them. We might cache the common description and photograph while fetching each customer's price. Choosing the reusable part is as useful as choosing the cache engine.
Sources and model notes
Working draft. Primary documentation accessed 2 October 2026. The product catalogue and logical clock are invented. Counts show work in this model, not measured product performance.
- Microsoft's cache-aside pattern: lookup, fill, and write ordering.
- Redis EXPIRE: key lifetime.
- Redis eviction policies: memory limits and approximate LRU.
- Redis persistence and Valkey persistence: snapshots, logs, and recovery.
- Valkey replication: asynchronous copies and the scope of WAIT.
The concurrent old-fill example follows from separating the catalogue read, catalogue write, deletion, and cache fill. The sequential experiment deliberately omits that race.