System design practice

Tiered CDN Cache

medium Based on a system published by Cloudflare caching cdn availability real-world

Cloudflare's Tiered Cache: edge misses fill from an upper tier, not the origin.

Solve it in your browser Read the lesson first

A CDN data center that does not have an asset asks the origin for it. With hundreds of data centers, each with its own cache, the origin sees the first request for a popular asset hundreds of times over, and its owner pays the cloud for every byte it sends out.

Cloudflare's Tiered Cache turns the flat network into a hierarchy. The data centers near visitors are the lower tier (the edge). When they miss, they ask an upper tier, a few data centers close to the origin, and only the upper tier asks the origin. Smart Tiered Cache picks the upper-tier data center for each origin and keeps a fallback in another location. Cloudflare says Tiered Cache cuts the cache miss rate by 60% or more.

Design the CDN side for one customer's site.

Functional requirements

Name the two tiers edge and upper: the tests in problem.proschi refer to them, and to the use case and scenario names above.

Scale

Constraints

What is given

problem.proschi declares the visitor, the owner and the origin, and holds the traffic, requirements and tests. Add the two tiers, the purge API, the connections and the two use cases.

The simulation is simpler than the real network: each tier is one node whose replicas stand for its data centers, the hit rates are the given mix rather than a result of the topology, and traffic between your own nodes is free, so the origin's egress bill shows up as load on the origin instead.

Based on

How your design is checked

You write the design as text in Proschi. Tests run in your browser: a simulation of the traffic above checks latency, availability, cost and what happens when a machine fails. How the simulation works.

Start designing

Lesson · 13 min read

Learn it: Tiered CDN Cache #

Tiered CDN Cache: one origin, hundreds of caches #

What you'll learn #

The problem, explained #

A CDN is a network of data centers that cache your content near visitors. When a data center does not have an asset, it fetches it from the origin, your own web servers. With hundreds of data centers each caching on its own, the origin sees the first request for a popular asset hundreds of times, and its owner pays the cloud for every byte it sends.

Cloudflare's Tiered Cache turns that flat network into a hierarchy. Data centers near visitors are the lower tier (the edge). On a miss they ask an upper tier: a few data centers close to the origin. Only the upper tier talks to the origin. You are designing the CDN side for one customer's site.

Two use cases:

Non-functional requirements: 50k requests per second, p99 under 150 ms, 99.99% availability, survival of any single machine (an upper-tier data center included), and $3,500 a month, the origin included.

What is given. given.proschi fixes the visitor, the owner and the origin: four servers that take 1k rps each, answer in 50 ms and cost $500 a month each. You cannot add origin servers; that is the point. A real origin is a customer's app in one cloud region, and the CDN's job is to protect it.

What the tests check: an edge miss asks the upper tier, and only an upper-tier miss reaches the origin; there is no path at all from the visitor or the edge to the origin; and a purge clears the upper tier before the edge and never touches the origin. The requirements add latency, availability, failure survival and cost.

The statement lists the model's simplifications: each tier is one node whose replicas stand for its data centers, hit rates are the given mix rather than an outcome of the topology, and traffic between your own nodes is free, so the origin's egress bill (what the cloud charges for data sent out) shows up as load on the origin instead.

Back-of-the-envelope #

Follow 50k requests a second down the hierarchy.

QuantityArithmeticResult
Requests at the edgegiven50k rps
Edge hits50k × 90%45k rps
Edge misses (to the upper tier)50k × 10%5k rps
Upper-tier hits5k × 60% (= 6% of all)3k rps
Origin fetches with an upper tier5k × 40% (= 4% of all)2k rps
Origin fetches without oneevery edge miss5k rps
Origin capacity4 × 1k rps4k rps
Purge fan-out50 purges × 300 edges15k deletes/s

The whole design is in two rows. With an upper tier the origin runs at 2k ÷ 4k = 50%. Without it, at 5k ÷ 4k = 125%: saturated. You cannot add origin servers, so the only lever is to send fewer requests there.

Why does an upper tier cut misses so much? A flat CDN with N data centers fetches each asset up to N times, once per cold cache. With an upper tier, each asset is fetched once by the upper tier, and the edges fill from it. Cloudflare reported a 60% or greater drop in the cache miss rate from tiering, which is the number this problem uses.

How the simulation sees it. The origin's utilisation at 50% with four servers gives a little queueing: its 50 ms base becomes a bit over 54 ms. At 125% the node is saturated and every latency requirement of a use case that loads it fails. The upper-tier miss path is edge (5 ms) + upper tier (5 ms) + origin (about 54 ms), around 64 ms on average; because only 4% of requests take it, the use case's p99 lands well under 150 ms. The flat design's p99, with a saturated origin, is several hundred milliseconds.

Cost. The origin is $2,000 of the $3,500, leaving $1,500. A CDN replica costs $100 in the simulation's price table and a service $100. CDN capacity (200k requests a second per replica) is far above what either tier sees, so replicas here are about failure, not load.

Failure. survive any node failure removes one replica of each node you run and re-runs the analysis. With one upper-tier data center, losing it breaks every edge miss, unless you add a fallback, which is where the classic mistake lives.

Concepts #

Cache hierarchies and origin shielding #

A cache hierarchy puts caches in layers: small caches near users, fewer bigger ones behind them, and the source of truth at the end. CPUs do it (L1, L2, L3, RAM), browsers do it (memory, disk, network), and CDNs do it with lower and upper tiers. Commercial CDNs call the upper tier an origin shield: one designated location that every miss for an origin must pass through.

Why it works: miss rates multiply. If the edge misses 10% and the upper tier misses 40% of what reaches it, the origin sees 10% × 40% = 4%. The upper tier also has a much better hit rate than any single edge, because it aggregates the misses of all edges: an asset requested once in Tokyo and once in Paris is two edge misses but only one upper-tier miss.

Trade-offs: an upper-tier hit adds a hop (often a long one, since the upper tier sits near the origin, not the visitor), so tiering slightly slows misses to make the origin's life much easier. It also concentrates load on a few data centers, so the upper tier must be redundant. Skip tiering when there is one cache location, when content is uncacheable (personalised pages), or when the origin is itself a scalable store such as S3 that does not care about load.

title "Two cache tiers"

user   "User"       [Actor]
near   "Near Cache" [CDN] x2
shield "Shield"     [CDN] x2
source "Source"     [Service] x2

user   -> near   : HTTPS
near   -> shield : fill
shield -> source : fill

usecase "Get" {
  user -> near : GET /a.png
  alt "Near hit" {
    near --> user : 200
  } alt "Shield hit" {
    near    -> shield : GET /a.png
    shield --> near   : 200
    near   --> user   : 200
  } alt "Miss" {
    near    -> shield : GET /a.png
    shield  -> source : GET /a.png
    source --> shield : 200
    shield --> near   : 200
    near   --> user   : 200
  }
}

Failover without stampedes #

When a cache layer fails, the tempting fallback is "go straight to the source". That is safe for a tiny cache and dangerous for a big one. The cache existed because the source cannot take the full load; the moment the cache disappears, the source receives everything it was shielded from, at once. This is a thundering herd: many clients missing on the same thing at the same time.

The right failover keeps the shape of the hierarchy: give the tier a second member of the same tier (a fallback upper-tier data center in another location) rather than a path that skips the tier. Cloudflare's later work on Smart Tiered Cache describes exactly this: a primary and a fallback upper tier for each origin. The origin can then also lock its firewall to the upper tier's addresses, which is why the tests forbid any path from the edge or the visitor to it.

Other tools in the same family: request collapsing (many concurrent misses for one key become one origin fetch), serving stale content while the origin is unreachable, and load shedding at the origin. Proschi models none of these directly, but the failure analysis makes the core lesson visible: a fallback that reaches a node which cannot take the load is not a fallback.

Purge ordering in a hierarchy #

Purging means deleting cached copies after the source changes. In a hierarchy, the order matters. Suppose you purge the edge first: an edge that receives a request between the two purges misses, asks the upper tier, which still has the old copy, and refills itself with it. The purge "succeeded" and the old asset is back.

The rule is purge from the source outward: the tier closest to the origin first, then the tiers that fill from it. Once the upper tier is clear, any edge miss falls through to the origin and gets the new asset. A purge never needs the origin itself, which already has the new version.

The fan-out is large: one purge must reach every edge data center. In Proschi an x300 prefix on a step means it happens 300 times per request; load counts all 300 calls while latency counts the step once, as if they ran in parallel. Alternatives to purging are versioned URLs (app.3f9a.js, never purged, just replaced) and short TTLs (time to live: how long a cached copy stays valid). Versioned URLs are the better default for static assets; purges are for content whose URL cannot change.

Designing it step by step #

1. Scope the problem #

Clarify: is the content cacheable and the same for everyone? (Yes: images and scripts.) How many edge locations? (About 300.) What is the origin's capacity, and can it grow? (4k rps, fixed.) How fast must a purge take effect? (Before the owner hears 200.) Fixed origin capacity is the constraint that drives everything; say so.

2. High-level design #

The starter is a flat CDN: visitor → edge → origin. Do the arithmetic from the table and show that the origin saturates at 125%. The obvious alternative, more origin servers, is ruled out by the constraints, and in reality it costs servers and egress that a better cache hit makes unnecessary.

The design that wins adds a node with id upper between edge and origin: visitor → edge → upper → origin. Write the three fetch scenarios as three branches of one alt. The edge-hit branch answers from the edge and touches nothing else. The upper-tier-hit branch adds one hop. Only the upper-tier-miss branch reaches the origin.

Then add a purge service for the owner: owner → purge API → upper tier, then edge.

3. Deep dive #

Redundancy at every tier. Every node you run needs to survive losing one replica. For the edge and the purge API that is just two replicas. For the upper tier, resist the fallback-to-origin idea from the concepts section and give the tier a second data center instead. Check the analysis: no single points of failure, and the origin still at the same utilisation after a failure.

The purge. Clear the upper tier first, then fan out to every edge with an x300 step, then answer the owner. Check the flow test that asks for upper before edge. Note what the purge does not do: it never calls the origin.

Latency. Look at the per-scenario numbers in the analysis. The miss path dominates p99 because of the origin's 50 ms, so keep the origin unsaturated and the p99 follows. If p99 is high, check the origin's utilisation before anything else.

Budget. Add up replicas times price, plus the $2,000 origin. The tiers are cheap compared with the origin you are protecting.

4. Wrap-up #

Summarise: two tiers, so 96% of requests never reach the origin; a redundant upper tier instead of an origin fallback; purges from the upper tier outward. Mention what the model leaves out: the real hit rates depend on topology and asset popularity, an upper tier far from the visitor adds latency to misses, and request collapsing at the upper tier would cut origin load further during a cold start.

Common mistakes #

Every edge fills from the origin (wrong/every-edge-fills-from-origin.proschi). The flat CDN: no upper tier, edges go straight to the origin. In the real world this is the default setup of most CDNs, and it works while the origin is big enough for the number of edges. When it is not, cold caches after a deploy or a purge overload it. Here it fails three ways: "Edge misses go to the upper tier, not the origin", "Only the upper tier talks to the origin", and p99, because 5k rps saturates a 4k-rps origin.

One upper-tier data center (wrong/single-upper-tier.proschi). Everything flows correctly, but the upper tier has a single replica and no fallback. When it fails, no edge miss can be filled. It fails survive any node failure. In production the same mistake looks like a single "shield POP" that becomes a global outage for cache misses when it goes down.

Upper tier down, so go to the origin (wrong/upper-down-goes-to-origin.proschi). A single upper tier with a -x failure scenario that falls back to the origin. It sounds resilient, but when the upper tier is down the origin takes every edge miss, the very overload the tier was there to prevent, and the origin can no longer restrict who talks to it. It fails "Only the upper tier talks to the origin", because the design now has an edge-to-origin connection.

Purging the edge first (wrong/purge-edge-first.proschi). Same design, purge order reversed. Between the two purges an edge can refill the stale copy from the upper tier, and the stale asset survives the purge. It fails "A purge clears the upper tier before the edge".

A load balancer as the edge. A load balancer distributes requests but caches nothing; every request reaches whatever is behind it. Use a CDN technology such as [Cloudflare] for both tiers.

In the interview #

Open with the observation that drives the design: "With hundreds of edge caches, the origin sees each asset's first request hundreds of times. I'd add a shield tier near the origin so it sees each asset roughly once." Then show the hit-rate arithmetic: 10% edge misses times 40% upper-tier misses equals 4% at the origin, half of its capacity.

Expected follow-ups:

Further reading #

Now design it

More system design problems