When a web page reloads in a fraction of the time it took the first time, the thing working behind the scenes is usually a cache. Caching keeps a copy of frequently requested data somewhere fast so the system avoids doing the same work over and over. This guide explains what a cache is, which layers it runs on, and what it actually speeds up when it is set up correctly in a publishing stack.
Related reading: SEO-friendly news software guide · How to build a fast database layer · How to optimize images on websites
What Is a Cache?
A cache is a technique for storing a copy of data that is expensive to compute or fetch in a temporary store that can be reached faster. The goal is simple: when the same request arrives a second time, serve the answer directly from that fast store instead of going back to the original source (a database, a disk, a remote server).
The idea repeats at every layer of a computer. A CPU uses small L1/L2/L3 caches that are faster than main memory; the operating system keeps disk blocks in RAM; the browser stores files it has downloaded; and a server holds a rendered HTML page ready instead of building it again. The underlying logic is always the same: produce once, use many times.
Why Caching Matters
The value of caching comes down to one sentence: it removes repeated work. Serving a ready-made copy instead of recomputing the same data or fetching it from disk on every request protects three core resources at once.
- Speed: Reading from memory takes far less time than hitting a disk or a remote database, which means a faster response for the user.
- Scalability: When most requests are served from the cache, the original source can carry far more traffic on the same hardware.
- Cost: Fewer database queries and fewer origin requests mean less CPU and less bandwidth consumed.
- Resilience: During sudden traffic spikes the cache acts as a buffer and protects the original source from being overloaded.
How Caching Works: Hits and Misses
Caching rests on two basic outcomes. When a request arrives, the system checks the cache first:
- Cache hit: The requested data is in the cache and still valid. The response comes straight from there; the original source is never touched. This is the fastest path.
- Cache miss: The data is not in the cache, or it has expired. The system goes to the original source, fetches the data, returns it to the user, and usually writes it into the cache for the next request.
The measure of how useful a cache is called the hit ratio: the share of total requests served from the cache. As the hit ratio rises, load on the original source falls and the average response time drops.
Cache space is not unlimited. When it fills up, an eviction policy decides what to drop to make room. The most common approaches are:
| Policy | Rule | When it fits |
|---|---|---|
| LRU | Evicts the least recently used entry | General purpose; when access has locality |
| LFU | Evicts the least frequently accessed entry | When certain records stay popular |
| FIFO | First in, first out | Simple; when access frequency does not matter |
| TTL | Marks an entry invalid once its time expires | When data going stale after a set time is acceptable |
The Layers of Caching
A single web request can be cached at several stops between the user and the database. Each layer solves a different problem; a strong setup uses them together.
| Layer | Where it runs | Typical use |
|---|---|---|
| Browser cache | On the visitor's device | Static files: CSS, JS, images, fonts |
| CDN / edge cache | On distributed servers near the user | Static assets and cacheable pages |
| Reverse proxy | In front of the application server | Full-page HTML cache (Varnish, Nginx) |
| Application / object cache | Next to the app (in RAM) | Query results, fragments, sessions (Redis, Memcached) |
| Database cache | Inside the database engine | Buffer pool, query plans, hot pages |
Browser Caching and HTTP Headers
The browser cache is governed by the HTTP response headers the server sends. Through these headers the server states how long a file may be stored and when it must be revalidated. The most important header is Cache-Control.
# Store for a year; ideal for version-stamped static files
Cache-Control: public, max-age=31536000, immutable
# Store, but revalidate with the server before each use
Cache-Control: no-cache
# Never store (sensitive, per-user responses)
Cache-Control: no-store
# Serve the stale copy, refresh in the background
Cache-Control: max-age=60, stale-while-revalidate=600Here is what the common directives mean:
| Directive | Meaning |
|---|---|
| max-age=N | The copy is considered fresh for N seconds |
| s-maxage=N | A separate lifetime for shared caches (CDNs) |
| public | May be cached by anyone |
| private | Only the browser stores it, not shared caches |
| no-cache | Stored, but revalidated before use |
| no-store | Not stored anywhere |
| immutable | Do not revalidate until it expires |
To check whether a copy is still valid, the server also sends validators: an ETag (a fingerprint of the content) and Last-Modified (the last change date). On the next request the browser sends these back via If-None-Match or If-Modified-Since; if the content has not changed, the server returns a short 304 Not Modified instead of the body, saving bandwidth.
CDN and Edge Caching
A CDN (Content Delivery Network) is a network that spreads copies of content across geographically distributed servers. The visitor gets the response from the nearest edge node rather than the origin server. This lowers latency and reduces load on the origin.
CDNs act on the s-maxage and Cache-Control headers. A CDN is ideal for static assets (images, CSS, JS). When you need to update cached copies by hand, you use the CDN's purge feature. To also make images lighter in size and format, see our image optimization guide.
Application and Object Caching
On the server side, object caches hold repeatedly computed data in memory. The two best known are Redis and Memcached; both store data as key-value pairs in RAM and respond very quickly because they read from memory rather than disk.
The most common pattern is cache-aside: the application checks the cache first, and if the data is not there it fetches it from the original source and writes it into the cache for later requests.
def get_article(article_id):
key = f"article:{article_id}"
data = cache.get(key) # 1) check the cache first
if data is not None:
return data # cache hit
data = db.query(article_id) # 2) cache miss -> original source
cache.set(key, data, ttl=300) # 3) store for 5 minutes
return dataObject caching is especially effective at reducing database load. Rather than running the same query thousands of times, storing the result for a while both relieves the database and speeds up the response. For designing the database layer to be fast from the start, our fast database layer guide is a good companion.
Database Caching
Database engines have their own caching mechanisms. The most important is the buffer pool (buffer cache): frequently read data and index pages are kept in RAM so the engine needs to hit the disk less often. Query plans and prepared statements are also cached, which avoids re-parsing the same query.
Cache Invalidation
The hardest part of caching is deciding when a copy has gone stale. If the original data changes but the old copy is still served, the user sees content that is out of date. The main ways to solve this are:
- Time-based (TTL): Each copy is given a lifetime and expires automatically. Simple, but it leaves a gap between an update and its visibility.
- Event-based purging: When data changes (for example, an article is updated) the relevant key is deleted or refreshed by hand. Fresher, but it requires more code.
- Version stamping (cache busting): A content hash is added to the file name (for example
style.a1b2c3.css); when the content changes the URL changes and the browser downloads the new file.
A Caching Strategy for News Sites
News sites are a special case for caching: traffic spikes suddenly (when a story goes viral), content is mostly read but rarely changed, and a headline can be updated within minutes. A balanced strategy combines these elements:
- Long-lived browser and CDN caching for static assets (images, CSS, JS), together with version stamping.
- Short-lived full-page caching for pages that change rarely (category, archive).
- Very short TTL or event-based purging for fast-changing areas such as the home page and headlines.
- Disabling the cache for personalized sections (logged-in users, comment forms), or assembling fragments at the edge.
On the home page, freshness comes first; on inner pages, speed does. A one-minute delay in updating a headline is usually acceptable; weigh that against the cost of rebuilding the whole page on every request.
Common Caching Mistakes
- Putting per-user data in a shared cache: one user's session could be served to another. Use
privateorno-storefor such responses. - Long max-age without version stamping: if you give a file a one-year lifetime and never change its name, your updates will not reach users.
- Forgetting invalidation: the content changes but the cache keeps serving the old copy.
- Caching everything: caching rarely requested or constantly changing data wastes memory and lowers the hit ratio.
Frequently Asked Questions
Are a cache and a CDN the same thing?
No. A cache is a general concept; a CDN is a service that applies that concept across geographically distributed servers. A CDN is one kind of caching layer, but not the only kind.
How long should I set the cache lifetime?
It depends on how often the data changes. You can give version-stamped static files a very long lifetime. Give frequently changing content a short TTL or use event-based purging. There is no single correct number; you decide the balance between freshness and speed.
Does no-cache really turn caching off?
No; the name is misleading. no-cache allows the copy to be stored but forces revalidation with the server before each use. For data you truly want stored nowhere, use no-store.
What is the difference between Redis and Memcached?
Both are fast in-memory key-value stores. Memcached is leaner and focused on simple caching scenarios; Redis offers richer data structures such as lists, sets and sorted sets, persistence options, and extra features like publish-subscribe. If plain caching is all you need, either works; if you need more data structures, Redis is more flexible.
What happens if cached data is lost from RAM?
Because a cache only holds copies, no data is lost. If the server restarts or the cache is cleared, the first incoming requests become cache misses and refill it from the original source. The system may run slowly for a short while; this is often called the cache "warming up".