Caching Strategies: Cache-Aside, Write-Through, and Write-Behind
A ticketing platform’s backend team once watched their primary database’s CPU pin at 100% the moment tickets for a popular concert went on sale. Every visitor’s page load triggered the same query, fetch the event details, the venue seating chart, and current availability, thousands of times a second, for data that changed maybe once every few seconds as tickets sold.
The database wasn’t struggling because the data was complex to fetch; it was struggling because it was answering the exact same question over and over for an audience that could have shared one answer.
Adding a cache in front of that query didn’t just make things faster, it was the difference between the sale succeeding and the database falling over before the first ticket sold.
Read Load That Took Down a Database at Peak Traffic
Caching exists to avoid repeating expensive work when the answer hasn’t changed since the last time someone asked. A cache sits between the application and a slower, more authoritative data source, usually a database, sometimes an external API or a computed result, and serves recent answers from fast, typically in-memory storage instead of recomputing or refetching them.
The strategy you choose determines when data gets written into the cache, how the cache stays synchronized with the source of truth, and what happens when they inevitably disagree for a moment. Three patterns dominate real-world systems: cache-aside, write-through, and write-behind, and each makes a different trade-off between consistency, latency, and implementation complexity.
- Cache hit: a request the cache can answer directly, without touching the underlying data source.
- Cache miss: a request the cache can’t answer, requiring a fetch from the source of truth.
- Staleness: the gap between what the cache holds and what the source of truth currently holds.
- Eviction: removing an entry from the cache, either because it’s outdated or because the cache needs room for something else.
Cache-Aside: Lazy Loading on Demand
Cache-aside, also called lazy loading, is the most common pattern because it requires the least upfront coordination between the cache and the data source.
The application checks the cache first; on a hit, it returns the cached value directly. On a miss, it fetches from the database, returns the result to the caller, and separately writes that result into the cache so the next request for the same key hits the cache instead.
def get_event_details(event_id):
cached = redis_client.get(f"event:{event_id}")
if cached is not None:
return json.loads(cached)
event = database.query_event(event_id)
redis_client.setex(f"event:{event_id}", 300, json.dumps(event))
return event
The application, not the cache itself, owns the logic for populating and invalidating entries, which gives it flexibility but also makes it responsible for getting that logic right.
A common failure mode is forgetting to invalidate or update the cache when the underlying data changes, an admin updates an event’s venue, but the cached copy keeps serving the old venue for up to five minutes until the TTL expires, which may or may not be acceptable depending on how time-sensitive that particular piece of data is.
- Simple mental model: read from cache, fall back to source on miss, write result back.
- Cache only stores what’s really requested: unlike some other patterns, nothing gets cached preemptively.
- Resilient to cache failures: if the cache is unavailable, the application can fall back to querying the source directly, just more slowly.
- Requires explicit invalidation logic: the application must actively clear or update cached entries when the source data changes.
Cache-aside also tends to be the easiest pattern to retrofit onto an existing system, since it doesn’t require touching every write path in the codebase, only the read path needs to change, wrapping an existing database call with a check-cache-first step.
This incremental adoption story is a large part of why teams reach for cache-aside first, even when a different pattern might technically suit a specific piece of data better: the cost of introducing it is low, and the benefit shows up immediately on the exact queries causing the most load, without a broader rewrite of how the application handles writes.
Write-Through Consistency Guarantees
Write-through caching flips the responsibility around for writes: every write to the data source also goes through the cache at the same time, in the same operation, so the cache is always updated as part of the write path rather than waiting for the next read to populate it.
def update_event_venue(event_id, new_venue):
database.update_event_venue(event_id, new_venue)
event = database.query_event(event_id)
redis_client.setex(f"event:{event_id}", 300, json.dumps(event))
This keeps the cache consistently up to date immediately after every write, eliminating the staleness window cache-aside can leave open. The cost is added latency on writes, since every write now waits on both the database update and the cache update to complete, and added complexity in ensuring the two stay in sync if one of those two operations fails partway through.
Write-through fits well for data that’s read far more often than it’s written and where staleness is truly unacceptable, a ticket inventory count during an active sale, for instance, where a stale cached “3 tickets remaining” could let more people believe they can still buy a ticket than really exist.
A subtler benefit of write-through is that it avoids the “thundering herd on first read” problem entirely for the data it covers, since the cache is always populated proactively rather than left empty until the first reader happens to request it. Cache-aside, by contrast, always has a cold-start moment for any given key, the very first request for a newly created event has to hit the database regardless of how well-tuned the caching layer is, because nothing has populated that key yet.
For workloads where even that first-request latency matters, write-through’s proactive population closes a gap cache-aside structurally cannot.
Write-Behind and Asynchronous Persistence
Write-behind (sometimes called write-back) caching takes the opposite approach on latency: writes go to the cache immediately and are acknowledged to the caller right away, while the actual write to the slower, durable data source happens asynchronously afterward, often batched with other pending writes for efficiency.
This delivers the lowest possible write latency, since the caller doesn’t wait on the database at all, and it can dramatically reduce database load for write-heavy workloads by batching many individual writes into fewer, larger database operations. The trade-off is a real risk of data loss: if the cache fails before the asynchronous write to the durable store completes, whatever hasn’t been persisted yet is gone,
unless the cache itself has been configured with its own durability guarantees.
- Lowest write latency: callers get an immediate acknowledgment without waiting on the durable store.
- Batching efficiency: many individual writes can be combined into fewer, more efficient database operations.
- Data loss risk on cache failure: unpersisted writes can be lost if the cache crashes before flushing to the durable store.
- Complex failure recovery: requires careful design to avoid silently losing writes during a cache outage or restart.
Write-behind tends to show up in high-throughput scenarios where losing a small amount of the most recent data during a rare failure is an acceptable trade-off for consistently low write latency, analytics event ingestion or activity logging are common candidates, where a financial ledger typically is not.
Implementations mitigate the data-loss risk in a few common ways rather than accepting it outright. Some pair the in-memory cache with a persistent, replicated write-ahead log of pending writes, so a cache restart can replay anything that hadn’t yet reached the durable store instead of losing it outright.
Others cap how long a write is allowed to sit unflushed, batching aggressively for throughput but flushing at least every few seconds, bounding the maximum possible loss window to something the business can reason about and accept, rather than leaving it open-ended and dependent on how the cache happens to be configured at any given moment.
Eviction Policies and TTL Design
A cache has finite capacity, and eviction policy determines what gets removed when that capacity is reached, or what expires on its own before then. Time-to-live (TTL) settings expire entries automatically after a fixed duration regardless of how often they’re accessed, bounding the maximum staleness a cache-aside pattern can produce even if invalidation logic has a bug somewhere.
- Least Recently Used (LRU): evicts whichever entry hasn’t been accessed in the longest time, a good general-purpose default.
- Least Frequently Used (LFU): evicts whichever entry has been accessed the fewest times, favoring entries with sustained popularity over recency.
- Time-to-live (TTL) expiration: entries expire automatically after a set duration, independent of access patterns, bounding maximum staleness.
- Random eviction: simpler to implement and surprisingly competitive in some workloads, though rarely the first choice for a general-purpose cache.
Choosing a TTL is itself a balancing act specific to each piece of data, too short, and the cache barely reduces load on the database because entries keep expiring before they’re reused; too long, and staleness becomes a real user-facing problem. Data that changes rarely, like a product category list, can tolerate a TTL measured in hours; data tied to active inventory during a flash sale might need a TTL measured in single-digit seconds, or a write-through pattern instead of relying on expiration at all.
Cache Invalidation Challenges
Phil Karlton’s famous remark that cache invalidation is one of the two hard problems in computer science holds up because the failure modes are truly subtle. Invalidating too aggressively defeats the purpose of caching, sending most requests back to the slow source anyway. Invalidating too conservatively, or forgetting an invalidation path entirely, leaves stale data being served with no clear signal that anything is wrong, since a stale cache doesn’t throw an error, it just quietly answers incorrectly.
Multi-key invalidation is an especially common trap: an event’s details might be cached under several different keys (by event ID, by venue, as part of a “featured events” list), and updating the event without invalidating every one of those related cache entries leaves some of them showing outdated data even though the “main” cache entry was correctly refreshed.
Systems handle this either by tracking which cache keys depend on a given piece of source data explicitly, or by leaning more heavily on short TTLs so any missed invalidation self-heals within a bounded, acceptable window rather than persisting indefinitely.
Distributed Caching With Redis and Memcached

At any substantial scale, caching moves out of a single application process’s memory and into a dedicated, shared caching layer that multiple application instances can read from and write to consistently. Redis and Memcached are the two dominant choices, and they differ substantially beyond just being “a fast key-value store.”
Memcached is a simpler, pure in-memory cache, optimized for straightforward key-value storage with multi-threaded performance that can be an advantage for very high-throughput, simple caching needs.
Redis offers a much richer feature set, data structures beyond plain strings (lists, sets, sorted sets, hashes), optional persistence to disk, built-in pub-sub, and Lua scripting for atomic multi-step operations, which is why Redis has become the default choice for most new caching layers, reserving Memcached for cases where its simpler, highly parallel model is specifically advantageous.
- Redis: rich data structures, optional persistence, pub-sub, broader feature set for general-purpose use.
- Memcached: simpler, multi-threaded, often faster for pure key-value workloads at very high concurrency.
- Cache clustering: both support distributing data across multiple nodes for capacity beyond a single machine’s memory.
- Client-side considerations: consistent hashing on the client (or a proxy layer) determines which node a given key lands on, and matters for even load distribution.
Stampedes, Thundering Herds, and Mitigations
A cache stampede happens when a popular cache entry expires and a burst of simultaneous requests all miss the cache at once, each independently going to the database to recompute the same value, momentarily creating exactly the kind of load spike the cache was supposed to prevent.
This is especially damaging for expensive-to-compute values with high concurrent access, precisely the ticketing platform’s scenario, where an expiring cache entry for a hot event could suddenly send thousands of simultaneous requests straight to the database.
- Request coalescing: letting only the first request that misses the cache really query the source, while concurrent requests for the same key wait for that result rather than each querying independently.
- Probabilistic early expiration: recomputing a cache entry slightly before its TTL really expires, staggered randomly across concurrent readers, to avoid many clients expiring at the exact same instant.
- Stale-while-revalidate: serving the slightly stale cached value immediately while asynchronously refreshing it in the background, avoiding a synchronous wait on the slow source entirely.
- Locking around cache population: using a short-lived lock so only one process repopulates a given key at a time, while others wait briefly or serve a stale fallback.
Final Thoughts
Caching strategies are, at their core, decisions about where staleness is allowed to live and for how long, since a cache that’s always perfectly consistent with its source isn’t really saving any work at all.
Cache-aside earns its popularity through simplicity and resilience, write-through trades write latency for immediate consistency, and write-behind trades a sliver of durability risk for the lowest possible write latency. None of the three is universally correct, and most real systems end up applying different strategies to different pieces of data based on how each one is really read and written.
The ticketing platform’s fix wasn’t picking the fanciest caching pattern available, it was recognizing that a short-TTL cache-aside layer in front of one specific, brutally repeated query was enough to turn a database-melting flash sale into an ordinary Tuesday.
Frequently Asked Questions
1. How do you decide between cache-aside and write-through for a given piece of data?
Cache-aside fits data that’s read much more often than written, where a brief staleness window after a write is tolerable. Write-through fits data where consistency immediately after a write matters more than write latency, such as inventory counts during high-contention events.
2. Can a single application use more than one caching strategy at once?
Yes, and most production systems do, applying cache-aside to most read-heavy data while reserving write-through or write-behind for specific fields or entities where their particular trade-offs really matter, rather than applying one strategy uniformly across an entire application.
3. What’s a reasonable default TTL when you’re not sure?
A short-to-moderate TTL, often measured in minutes, is a safer starting point than an aggressive one measured in hours, since it bounds worst-case staleness while still substantially reducing database load; it can be tuned upward once real traffic patterns and staleness tolerance are better known.
4. Does caching help with write-heavy workloads?
Less directly than with reads, though write-behind caching specifically targets write latency by deferring the durable write. For most systems, caching’s primary benefit remains reducing redundant read load rather than accelerating writes.
5. How do you monitor whether a cache is really helping?
Track the hit rate (the percentage of requests served from cache rather than falling through to the source) alongside the latency and load on the underlying data source. A low hit rate suggests either a TTL that’s too short, a cache that’s too small, or an access pattern with too little repetition to benefit from caching in the first place.
6. What happens if the cache and the database disagree?
The database is treated as the source of truth in virtually all caching strategies, so a disagreement means the cache is stale and needs to be invalidated or refreshed, not the other way around; well-designed systems bound how long that disagreement can persist through TTLs or explicit invalidation logic.
