On this page
The simulation behind Analysis, Tests and practice verdicts
How the simulation works, and how far to trust it.
Every number in the editor's Analysis and Tests tabs and every practice
verdict comes from one small analytical model: a handful of formulas over your diagram, its
traffic and a table of default numbers per technology. It runs in milliseconds, has no
randomness and never measures anything. Same text in, same numbers out.
That makes it good at one thing: telling a sound design from a broken one, and showing which node gives out first. It is not a load test and not a capacity plan. This page lists every rule the model uses, with the real default numbers, then everything it leaves out and which way each omission pushes the numbers.
01
What is modelled
The code is in frontend/src/sim/
(profiles.ts, analyze.ts, flow.ts, tests.ts). Each rule
below is what that code does today; the numbers in this page are checked against it by a test.
-
1.1 Load comes from traffic
A use case with rate R splits its traffic over its scenarios by
mix; withoutmixit all goes to the first scenario, and shares that do not add up to 100% are scaled. Every request step that gets through (->and->>) adds its rate to the node it targets. Responses add nothing, and neither does a failed call (-x): in that scenario its target is down, so it receives nothing and sends nothing back.Every step counts, also steps after the response and the work an async consumer does later, so a queue's consumer carries load even though the user never waits for it. A use case without a traffic line adds no load.
load(n) = Σ R(U) × share(s) × N(step) over every request step to n that is not -x in every scenario s of every use case N(step) = the x<N> fan-out, 1 without one -
1.2 Reads and writes
Each request step is a write when its HTTP method is POST, PUT, PATCH or DELETE, or, without a method, when the first word of its label is a write verb: INSERT, UPDATE, UPSERT, DELETE, SET, PUBLISH, SEND, ENQUEUE, CHARGE, CREATE, SAVE, STORE, UPLOAD, RESERVE, HOLD and about thirty more (full list). Everything else (GET, SELECT, QUERY, SCAN, LOOKUP, an event name) is a read. A request to a queue is always a write, whatever its label.
Reads and writes are counted separately against separate capacities, and only writes count for
durableandwrites X before responding.- Write
api -> db : INSERT order- Write
user -> api : POST /orders- Write
api ->> jobs : OrderPlacedqueue- Read
api -> cache : GET feed:{user}- Read
api -> db : SELECT posts
-
1.3 Replicas, shards and single-primary writes
A node declared with
x<n>has n replicas;capacity { db shards 4 }splits a store into independent partitions, each with its own replicas. Every replica of every shard is assumed to get an equal slice of the load.Relational databases (PostgreSQL, MySQL, Aurora, RDS, SQL Server, …) are single-primary: their replicas serve reads, but every write goes to the one primary of its shard, so replicas add read capacity and only shards add write capacity. Everything else, including partitioned NoSQL stores, caches and queues, serves reads and writes on every replica.
A node's request utilisation is the busier side for a single-primary store (replicas and primaries work in parallel) and the sum of both shares for everything else (they share the same machines). Its utilisation is the larger of that and its bandwidth utilisation (1.6). At or above 100% the node is saturated: every latency requirement of a use case that sends it load fails. Above 70% it is shown as hot.
read capacity = reads × replicas × shards write capacity = writes × shards single-primary = writes × replicas × shards everything else ρ_read = read load ÷ read capacity ρ_write = write load ÷ write capacity ρ_req = max(ρ_read, ρ_write) single-primary = ρ_read + ρ_write everything else ρ = max(ρ_req, ρ_bandwidth) -
1.4 Utilisation turns into queueing
Each node has a base latency per call when idle (from its profile, below). Under load a request may wait for a free server. The model treats each replica of a shard as one server whose service time is that base latency, and uses the M/M/c queue: with c servers at utilisation ρ, a request waits with the Erlang C probability C, on average
base ÷ (c(1 − ρ)). Load is spread evenly over shards and a key's requests can only go to its own shard, so c is one shard's replicas; writes that bind a single-primary store queue for its one primary.One server gives the classic
1 ÷ (1 − ρ)slowdown. A pool waits much less at the same utilisation: at 50%, one server waits half the time, four servers 17%. Utilisation is capped at 95% in this formula, so a saturated node has a finite latency (at most 20× its base) and fails its requirements through saturation instead.C = Erlang C(c, ρ) ρ capped at 0.95 hop(n) = base + C × base ÷ (c (1 − ρ)) = base ÷ (1 − ρ) for c = 1Slowdown over the idle latency by utilisation ρ 10% 50% 70% 90% ≥ 95% 1 server 1.1× 2× 3.3× 10× 20× 4 servers 1× 1.09× 1.36× 2.97× 5.46× -
1.5 Latency and percentiles
A scenario's mean latency is the sum of the hops on the synchronous critical path of its entry request: requests sent before the entry request is answered, by the entry's sender or by a node still working on it synchronously. A
pargroup counts its slowest member; an async send (->>) counts its own hop but not the receiver's work; a failed call (-x) costs a fixed timeout of 1000 ms instead of the target's latency, or whatcapacity { db timeout 200ms }sets for that node.Percentiles are computed, not measured. Each hop's time is a fixed part (half its service time) plus an exponential tail that carries the rest of its mean: the other half of the service time and all of the queueing. An idle hop's p99 is 2.8× its mean; one server at 50% makes it 3.7×, at 90% 4.4×. Timeouts and transfer time are fixed: they add once to every percentile. A path takes every hop at the same quantile, which errs on the slow side for long paths.
A use case's percentile is the percentile of all its requests together: the scenarios mixed by their shares. With 10% cache misses, p99 is close to the miss path's p90; it moves smoothly as a share grows, without jumps. Both are solved numerically, the same way every time.
item = fixed + Exp(tail) fixed = base ÷ 2 + transfer (a -x: the timeout, no tail) tail = hop(target) − base ÷ 2 p_q(s) = Σ over the parts of the critical path of max over the items of the part of fixed + tail × ln(1 ÷ (1 − q)) p_q(U) = the t with Σ share(s) × P(s ≤ t) = qPercentiles of an idle hop over its mean idle hop p50 p90 p95 p99 p999 ÷ mean 0.85× 1.65× 2× 2.8× 3.95× -
1.6 Fan-out, payload size, bandwidth and egress
A label that starts with
x200means the step happens 200 times per request: load counts 200 calls; latency counts the step once (as if the calls were batched or parallel), but its 200 payloads all cross the link.A label with
~2MBgives the step a payload: a read's payload travels back from the node it asks, a write's goes from the sender. It adds transfer time at the slower of the two ends' per-replica bandwidth (client 10 MB/s, edge, CDN, load balancer and gateway 1,000 MB/s, service 200 MB/s, everything else 100 MB/s). The bytes also fill the bandwidth of the nodes you run: a service whose payloads exceed its replicas' bandwidth is saturated. Object storage and CDNs scale out behind one name, so for them bandwidth is only the speed of one transfer.Payloads sent to a client or a third party leave your network and are charged as internet egress by the node that sends them: $0.09/GB from anything you run, $0.02/GB from a CDN. Traffic between your own nodes (storage to a service, a CDN filling from storage) is free, and what clients and third parties send is not your bill.
transfer(step) = N × size ÷ min(bandwidth(from), bandwidth(to)) ρ_bandwidth(n) = payload bytes in and out per second ÷ (bandwidth × replicas × shards) egress GB/month = R × share × N × size × 2,592,000 s ÷ 10⁹ for payloads sent to a client or a third party egress cost = egress GB × price of the node that sends it -
1.7 Availability
Each replica is up a fixed fraction of the time (profile below), independently of every other. A node is up when any of its replicas is. A use case is up when every node on the synchronous path of its main scenario (the one with the largest share) is up, except that a node with a fallback only takes the use case down when its fallback path is down too.
A write to a single-primary store needs its primary. With a second replica, failover promotes one, but writes wait out the promotion: the model keeps 10% of the primary's downtime, so a PostgreSQL primary with replicas takes writes 99.995% of the time, alone 99.95%. Reads still count every replica.
A fallback is a success scenario that calls the node with
-xbefore the entry response and completes without a successful synchronous call to it. Load and saturation do not affect availability.A(n) = 1 − (1 − a(n))^replicas A_write(n) = 1 − (1 − a(n)) × 0.1 single-primary, 2+ replicas = a(n) single-primary, 1 replica A(U) = Π over nodes n on the main path of A(n), or A_write(n) when the main path writes to it 1 − (1 − A(n)) × (1 − A(fallback)) with a fallback -
1.8 Cost
Each replica of each shard costs a flat monthly price from its profile, whatever its load, plus the egress above, for a 30-day month. Clients, DNS and external systems cost nothing. There are no instance sizes, tiers, reservations or per-request prices.
cost = Σ over nodes replicas × shards × cost(n) + egress cost(n) -
1.9 Failures: survive, fallbacks, single points of failure
survive any node failuretakes every node you run (not clients, external systems or DNS) and removes one instance. With two or more replicas, the whole design is analysed again with one replica fewer: the node must not saturate, and every latency requirement that held must still hold. On a sharded store the instance comes out of one shard; keys cannot move to another shard, so that shard's numbers are what count. A single-primary store that loses its primary promotes a replica: writes keep a primary and reads lose one replica, the same as losing a replica. With one replica, every use case whose success scenarios need the node must have a fallback scenario for it.survive failure of <selector>does the same for the selected nodes.The Analysis tab lists as single points of failure the single-replica nodes that some use case needs and has no fallback for. A failed call (
-x) in a scenario is how a design says "and this is what happens when that node is down".- x2+
- with one replica fewer, the node does not saturate and the latency limits still hold
- x1
- every use case that needs it has a fallback scenario
- Skipped
- clients, external systems, DNS, notes
-
1.10 Tech stacks, kinds and consistency
A node's tech stack picks its kind and its profile. The catalog knows 205 techs with the names people write as aliases (
[S3],[Postgres],[ALB],[k8s]), ignoring case, punctuation and a trailing version. A tech it does not know is a warning, and the node is simulated as the kind its name suggests ([TigerBeetle DB]is a database), else the kind of the closest catalog tech, else a service: never a free, infinitely fast client.The edge splits into
cdn,loadbalancer,gatewayanddns, each with its own numbers; a firewall ([WAF]) is a plainedge, andany edgeselects all of them. DNS is off the request path: it adds no latency, never fails and costs nothing. Every data store has a consistency,strongoreventual, used only by the selectorsany strong storeandany eventual storein tests; replication lag and staleness are not simulated.[Postgress] did you mean PostgreSQL? → database [TigerBeetle DB] unknown → database (from its name) [Cobol Mainframe] unknown → service any edge edge, cdn, loadbalancer, gateway, dns any strong store relational, MongoDB, Spanner, queues … any eventual store caches, DynamoDB, Cassandra, search, storage …
Default profiles, per replica
Teaching values, right to an order of magnitude, not benchmarks. capacity { … } overrides any
of them per node (see How to read the results).
| Kind | Examples | Reads/s | Writes/s | Latency | Avail. | Cost/mo | Bandwidth | Egress | Durable | Consistency |
|---|---|---|---|---|---|---|---|---|---|---|
| client | Actor, shapes without a tech | ∞ | ∞ | 0 ms | 100% | $0 | 10 MB/s | — | no | — |
| cdn | CloudFront, Azure CDN, Cloud CDN, Front Door | 200k | 200k | 5 ms | 99.99% | $100 | 1,000 MB/s | $0.02/GB | no | — |
| edge | WAF, AWS WAF, Global Accelerator | 100k | 100k | 2 ms | 99.99% | $50 | 1,000 MB/s | $0.09/GB | no | — |
| loadbalancer | AWS Load Balancer (ALB, NLB), nginx, Envoy, HAProxy | 100k | 100k | 2 ms | 99.99% | $50 | 1,000 MB/s | $0.09/GB | no | — |
| gateway | AWS API Gateway, Azure API Management, Kong | 10k | 10k | 10 ms | 99.95% | $100 | 1,000 MB/s | $0.09/GB | no | — |
| dns | Route53, Azure DNS, Cloud DNS | ∞ | ∞ | 0 ms | 100% | $0 | 1,000 MB/s | — | no | — |
| service | Service, REST API, gRPC, Spring Boot, Go, Node.js, ECS, EKS, Kubernetes | 2k | 2k | 10 ms | 99.5% | $100 | 200 MB/s | $0.09/GB | no | — |
| function | Lambda, Cloud Functions, Cloud Run | 10k | 10k | 25 ms | 99.95% | $200 | 100 MB/s | $0.09/GB | no | — |
| cache | Redis, Valkey, ElastiCache, Memcached | 100k | 100k | 1 ms | 99.9% | $150 | 100 MB/s | $0.09/GB | no | eventual |
| database, relational | PostgreSQL, MySQL, Aurora, RDS, SQL Server, Oracle | 20k | 5k | 5 ms | 99.95% | $400 | 100 MB/s | $0.09/GB | yes | strong |
| database, NoSQL | DynamoDB, Cassandra, ScyllaDB, Cosmos DB | 20k | 20k | 5 ms | 99.99% | $500 | 100 MB/s | $0.09/GB | yes | eventual |
| database, NoSQL | MongoDB, CockroachDB, Bigtable, Spanner | 20k | 20k | 5 ms | 99.99% | $500 | 100 MB/s | $0.09/GB | yes | strong |
| search | Elasticsearch, OpenSearch, Solr | 3k | 3k | 15 ms | 99.9% | $400 | 100 MB/s | $0.09/GB | yes | eventual |
| analytics | BigQuery, Snowflake, Redshift, ClickHouse | 200 | 200 | 500 ms | 99.9% | $300 | 100 MB/s | $0.09/GB | yes | eventual |
| queue | Kafka, Kinesis, SQS, SNS, Pub/Sub, RabbitMQ | 50k | 50k | 5 ms | 99.99% | $200 | 100 MB/s | $0.09/GB | yes | strong |
| storage | S3, GCS, Azure Blob, MinIO | 5k | 5k | 30 ms | 99.99% | $50 | 100 MB/s | $0.09/GB | yes | eventual |
| external | Stripe, Twilio, SendGrid, third-party APIs | 1k | 1k | 200 ms | 99.9% | $0 | 100 MB/s | — | no | — |
Relational databases take 5k writes per second per
shard (on the primary), not per replica. Egress is the price of data a node sends to clients and
third parties. A tech stack the catalog does not know gets the row of the kind its name suggests (a service
when nothing does) and a warning that names the closest known tech; to keep a product name the catalog
lacks without the warning, use the generic tech of its kind ([Database],
[Message Queue]) and put the product in the node's name.
02
Worked examples
Every number in this section comes from running the simulation on the source shown next to it, and a test fails if the code and the page disagree. Paste a source into the editor to see the same numbers in its Analysis and Tests tabs.
2.1 A cached read path
title "Feed"
user "User" [Actor]
api "Feed API" [REST API] x6
cache "Feed cache" [Redis] x2
db "Feed DB" [PostgreSQL] x2
user -> api : HTTPS
api -> cache : RESP
api -> db : SQL
usecase "Read feed" {
user -> api : GET /feed
alt "Cache hit" {
api -> cache : GET feed:{user}
api --> user : 200
} alt "Cache miss" {
api -> cache : GET feed:{user}
api -> db : SELECT posts
api -> cache : SET feed:{user}
api --> user : 200
}
}
usecase "Post" {
user -> api : POST /posts
api -> db : INSERT post
api --> user : 201
}
traffic {
"Read feed" 6k rps mix "Cache hit" 90%, "Cache miss" 10%
"Post" 1k rps
}
requirements {
p99 "Read feed" < 60ms
availability "Read feed" >= 99.9%
durable "Post"
survive any node failure
cost <= 2500 usd/month
}
Load. api receives every entry request: 6k reads and 1k writes a second
against 12k rps (6 × 2k), so
ρ = 50% + 8%
= 58%. db gets only the misses
(10% of 6k = 600 rps SELECTs) and every post
(1k rps INSERTs). Its two replicas serve
40k rps of reads but its one primary takes
5k rps of writes, so writes set its utilisation:
20%.
Latency. The six API replicas at 58% queue as one pool of 6 servers: a request waits only 18% of the time, so the API's 10 ms becomes 10.7 ms (one server at 58% would take 24 ms). The cache stays at 1 ms and the database at 6.3 ms; its writes queue for the one primary. The hit path is api + cache, mean 11.7 ms; the miss path adds the SELECT and the SET, mean 19 ms.
Percentiles. p99 is taken over all reads, 90% hits and 10% misses mixed: 38.6 ms, above the hit path's own p99 (34.1 ms) and below the miss path's (56.7 ms). p50 is 10.3 ms, a little under the hit path's mean: half of each service time is fixed, and the median of the exponential rest is below its mean.
Cost. 6 × $100 + 2 × $150 + 2 × $400 = $1,700 a month.
- pass
- p99 of Read feed is 38.6 ms (limit 60 ms)
- pass
- availability of Read feed is 99.9999% (limit 99.9%)
- pass
- "Post" writes to a durable store before responding
- pass
- Every use case keeps working after losing any of 3 nodes
- pass
- Total cost is $1,700/month (limit $2,500/month)
2.2 Replicas do not scale writes
title "Orders"
client "Client" [Actor]
api "Order API" [REST API] x8
db "Orders DB" [PostgreSQL] x3
client -> api : HTTPS
api -> db : SQL
usecase "Place order" {
client -> api : POST /orders
api -> db : INSERT order
api --> client : 201
}
traffic {
"Place order" 6k rps
}
requirements {
p99 "Place order" < 100ms
}
Then add two shards:
capacity {
db shards 2
}
Saturated. 6k INSERTs a second reach a PostgreSQL primary that takes 5k. Three replicas give 60k rps of read capacity, none of it usable for writes, so the database runs at 120% and the requirement fails however much time is left:
- fail
- p99 of Place order: db is saturated (reads 0 rps of 60k rps, 0%; writes 6k rps of 5k rps, 120%) (limit 100 ms)Writes to db (PostgreSQL) need 6k rps and each primary takes 5k rps; read replicas do not add write capacity. Add shards (capacity { db shards 2 } keeps writes under 70%), move the data to a partitioned store (DynamoDB, Cassandra), or batch writes through a queue
Sharded. With shards 2 there are two primaries, each with three replicas:
write capacity 10k rps, utilisation
60%, and the database hop takes
12.5 ms. The cost doubles too: 3 replicas × 2 shards ×
$400 = $2,400.
- pass
- p99 of Place order is 76.7 ms (limit 100 ms)
Moving the table to a partitioned store such as DynamoDB, which scales writes with replicas, would
pass as well. In a practice problem shards is the one capacity setting you may write
yourself, because how many shards to run is a design decision.
2.3 Bytes, bandwidth and egress
title "Photos through the API"
viewer "Viewer" [Actor]
api "Photo API" [REST API] x2
blobs "Photos" [AWS S3] x2
usecase "View photo" {
viewer -> api : ~2MB GET /photos/{id}
api -> blobs : ~2MB GET photo
api --> viewer : 200
}
traffic {
"View photo" 20 rps
}
title "Photos through a CDN"
viewer "Viewer" [Actor]
cdn "CDN" [AWS CloudFront] x2
blobs "Photos" [AWS S3] x2
usecase "View photo" {
viewer -> cdn : ~2MB GET /photos/{id}
alt "Edge hit" {
cdn --> viewer : 200
} alt "Edge miss" {
cdn -> blobs : ~2MB GET photo
cdn --> viewer : 200
}
}
traffic {
"View photo" 20 rps mix "Edge hit" 95%, "Edge miss" 5%
}
Transfer time. A 2 MB answer to the viewer moves at the viewer's 10 MB/s: 200 ms, more than every node's own latency together, and a fixed cost that adds once to every percentile. Through the API the photo also crosses api ↔ storage at storage's 100 MB/s (20 ms): mean 260 ms, p50 254 ms and p99 334 ms. Through the CDN, hits take 205 ms and misses 255 ms; with 5% misses, p99 is 266 ms. Latency barely improves, because the viewer's own bandwidth dominates both designs.
Egress. 20 photos of 2 MB a second is 103,680 GB a month sent to viewers. The API sends them at the internet rate: $9,331 of egress and $9,631 in total; what storage sends the API stays inside and costs $0. Through the CDN, the edge sends the same bytes for $2,074, and filling it from storage is free: $2,374 in total.
03
Simplifications that may surprise you
These follow directly from the formulas above. None of them changes which of two designs is better in ordinary cases, but each one moves absolute numbers in a known direction.
| Slow hops come together pessimistic | A path's percentile takes every hop at the same quantile, as if the slow requests of every node hit the same request. Real hops vary independently, so long paths have narrower tails than the model's. Kept because it is simple, deterministic and errs on the safe side; the sum of independent hops has no closed form. |
|---|---|
| A replica is one server either way | Queueing treats each replica as one server whose service time is its base latency. A real replica serves many requests at once, so its latency stays flat longer and then climbs faster. The shape (more replicas queue less, one busy server queues a lot) is right; the exact curve is a teaching value. |
| Fan-out latency counted once optimistic |
x200 multiplies load and transfer, but adds one hop of latency. The slowest of 200
parallel calls is far slower than a typical one (tail at scale). Kept, because the same label also
means a batched call, which really is one round trip.
|
| Independent replicas optimistic |
1 − (1 − a)^n assumes replicas never fail together: six service replicas at 99.5% each
show as 99.999999%. A bad deploy, a zone outage or a shared dependency
takes them all at once.
|
| Failover is a fixed share either way | Writes to a single-primary store with replicas lose 10% of the primary's downtime to failover, whatever the database or its settings. Real failovers take from seconds to minutes, and some lose the last writes. |
| Saturation does not spread optimistic | A saturated node still passes its full load downstream and its latency stops growing at 20×. Real queues grow without bound, time out and push retries upstream. Kept on purpose: a design with a saturated node already fails, and throttling the load would hide the next bottleneck behind the first, so you would fix them one at a time. |
| Managed bandwidth is unlimited optimistic | Object storage and CDNs never saturate on bytes; their bandwidth only sets the speed of one transfer. That is close to how they behave, but per-prefix request limits and origin shields are not modelled. |
04
What is not modelled
Most omissions make the model optimistic: real systems are slower, less available and dearer than the numbers say. Treat a design that only just passes as one that would fail.
| Cold starts optimistic | A function always takes its 25 ms. Real cold starts add hundreds of milliseconds to the tail, more for big runtimes. |
|---|---|
| GC pauses, CPU steal, noisy neighbours optimistic | Every replica is equally fast all the time. Pauses and shared hardware fatten p99 and p999 well beyond the tails the model computes from queueing alone. |
| Network latency and jitter optimistic | A hop costs the target's latency only; there is no separate round trip, no cross-zone or cross-region distance and no packet loss. Multi-region designs look as fast as single-zone ones. |
| Retries and thundering herds optimistic | A failed call costs one timeout and the scenario moves on. Retries that double the load on a sick node, retry storms after an outage and a cold cache stampeding the database do not exist. |
| Partial and grey failures optimistic | Nodes are up or down. A slow replica, a half-broken network path or a dependency that answers with errors at 2% is not representable, except as a scenario you write. |
| Cache hit rate either way |
The hit rate is whatever mix says. Warm-up, eviction, a cache emptied by a restart, and
hit rates that change with traffic or data size are not modelled. Pick the rate you would actually
measure, and model the cold case as its own scenario.
|
| Hot keys and skew optimistic | Load is spread evenly over every replica and shard. One celebrity account or one hot partition saturates a real shard long before the average does. |
| Traffic shape and autoscaling optimistic | Traffic is one constant rate and replica counts are fixed. Peaks, daily cycles, bursts and the minutes an autoscaler needs to catch up are not modelled. Enter your peak rate, not your average. |
| Contention and per-request work either way | Every request to a node costs the same capacity: a key lookup and a five-table join, a 1 KB and a 1 MB write. Locks, connection pools, transactions and replication work on followers are not modelled. |
| Queues and consumers either way | Backlog, consumer lag, ordering, redelivery and dead letters are out of scope. A queue is a node that takes writes; its consumers are steps with load. |
| Time and consistency either way | TTLs, expiry, replication lag and stale reads are not simulated; strong and eventual are labels for tests. Model expiry or a stale read as a scenario. |
| Third-party limits optimistic | External systems are free, take 1k requests a second per replica and never rate-limit you. |
| Real cloud pricing either way | One flat monthly price per replica. No instance sizes, reserved or spot pricing, per-request pricing (Lambda, DynamoDB on demand, API Gateway, S3 requests), storage at rest, free tiers or egress tiers. Traffic between your own nodes is free, also across zones, where clouds charge about $0.01/GB each way (optimistic, and small next to internet egress). |
05
How to read the results
-
5.1 Compare designs, find the bottleneck
The model is deterministic and its inputs are rough, so read its numbers relatively. Two designs under the same traffic and profiles compare fairly: a cache that takes p99 from 300 ms to 40 ms, a database that goes from 120% to 60%, a CDN that cuts egress by four. A 10% difference between two designs is inside the error of the teaching values.
The utilisation column shows where the design gives out first. Raise the traffic until something turns amber (above 70%) or red (saturated): that node, and the lever the hint names (replicas, shards, a cache, async work), is what the design is really about.
Do not size a fleet from these numbers or quote them as SLOs. A capacity plan needs load tests of your own services, your peak traffic and your cloud bill.
- Good for
- comparing designs, finding the first bottleneck, spotting single points of failure, sanity-checking a budget
- Not for
- capacity plans, SLO commitments, cost forecasts, tail-latency estimates
- Fixed
- no randomness, no warm-up: the same text always gives the same numbers
-
5.2 Calibrate with your own measurements
capacity { … }replaces any default for one node, per replica. Once you have real numbers, put them in and the rest of the model works from them:- Rate (
3k rps, orreads … writes …): the throughput at which one replica stops keeping up in a load test, since the model adds queueing delay on top. Not the rate at which latency is still fine. - Latency: the median per call at low load, measured from the caller, so it includes the network.
- Availability: per replica, from your provider's SLA or your incident history.
- Cost: what one replica costs on your bill. Egress and bandwidth: your prices and your instance's network.
- Shards, durable, consistency: how the store is actually set up.
Calibrated rates fix the capacity side; the omissions in section 04 still apply, so keep headroom.
capacity { api 3k rps latency 12ms cost 140 usd/month db reads 25k rps writes 6k rps availability 99.95% users shards 4 consistency strong blobs bandwidth 500 MB/s egress 0.05 usd/GB } - Rate (
06
Practice verdicts
-
6.1 What "Solved" means
A practice problem is solved when every requirement and every test in the problem's given file passes against your design, with the same simulation as the editor. The problem fixes the traffic, the limits and any capacity numbers; a
capacityline in your own file is an error, exceptshards <n>, so you cannot buy your way out with a bigger node.Passing means the design is sound under this model: the requests flow the right way, nothing saturates at the given traffic with the default numbers, every node you run can fail without breaking a use case, and the bill fits the budget. It does not mean the system would hold up in production; section 04 is still true.
- Latency
- p99 limits fail when the path is too long, a node queues, or any node a use case loads is saturated
- Failure
- survive and availability fail on single points of failure
- Cost
- the budget fails designs that throw replicas at the problem
- Flow
- tests check the key idea: cache before database, write before responding, never wait for a queue
-
6.2 Why brute force fails
Each problem's limits are set between the reference solution and the shortcuts a candidate might try: loose enough that the intended design passes with room, tight enough that the shortcut does not. More replicas fix saturation but break the cost budget; skipping the cache saves its cost but loads the database and the p99; a load balancer in front of storage fails the flow test that asks for a CDN.
Every problem keeps those shortcuts as
wrong/*.proschidesigns, each naming the tests it must fail.proschi problem checkruns them in CI, so a change to the model or the limits that lets a shortcut pass is caught before it ships. How problems are calibrated.# expect-fail: p99 of Hold seat < 130 ms # Scales the relational database with read replicas instead of shards. import "problem.proschi" # … the rest of the design … db "Seats DB" [PostgreSQL] x4