Docs menu
On this page

The simulation behind Analysis, Tests and practice verdicts

How the simulation works, and how far to trust it.

Every number in the editor's Analysis and Tests tabs and every practice verdict comes from one small analytical model: a handful of formulas over your diagram, its traffic and a table of default numbers per technology. It runs in milliseconds, has no randomness and never measures anything. Same text in, same numbers out.

That makes it good at one thing: telling a sound design from a broken one, and showing which node gives out first. It is not a load test and not a capacity plan. This page lists every rule the model uses, with the real default numbers, then everything it leaves out and which way each omission pushes the numbers.

01

What is modelled

The code is in frontend/src/sim/ (profiles.ts, analyze.ts, flow.ts, tests.ts). Each rule below is what that code does today; the numbers in this page are checked against it by a test.

  1. 1.1 Load comes from traffic

    A use case with rate R splits its traffic over its scenarios by mix; without mix it all goes to the first scenario, and shares that do not add up to 100% are scaled. Every request step that gets through (-> and ->>) adds its rate to the node it targets. Responses add nothing, and neither does a failed call (-x): in that scenario its target is down, so it receives nothing and sends nothing back.

    Every step counts, also steps after the response and the work an async consumer does later, so a queue's consumer carries load even though the user never waits for it. A use case without a traffic line adds no load.

    load(n) = Σ  R(U) × share(s) × N(step)
              over every request step to n that is not -x
              in every scenario s of every use case
    
    N(step) = the x<N> fan-out, 1 without one
  2. 1.2 Reads and writes

    Each request step is a write when its HTTP method is POST, PUT, PATCH or DELETE, or, without a method, when the first word of its label is a write verb: INSERT, UPDATE, UPSERT, DELETE, SET, PUBLISH, SEND, ENQUEUE, CHARGE, CREATE, SAVE, STORE, UPLOAD, RESERVE, HOLD and about thirty more (full list). Everything else (GET, SELECT, QUERY, SCAN, LOOKUP, an event name) is a read. A request to a queue is always a write, whatever its label.

    Reads and writes are counted separately against separate capacities, and only writes count for durable and writes X before responding.

    Write
    api -> db : INSERT order
    Write
    user -> api : POST /orders
    Write
    api ->> jobs : OrderPlaced queue
    Read
    api -> cache : GET feed:{user}
    Read
    api -> db : SELECT posts
  3. 1.3 Replicas, shards and single-primary writes

    A node declared with x<n> has n replicas; capacity { db shards 4 } splits a store into independent partitions, each with its own replicas. Every replica of every shard is assumed to get an equal slice of the load.

    Relational databases (PostgreSQL, MySQL, Aurora, RDS, SQL Server, …) are single-primary: their replicas serve reads, but every write goes to the one primary of its shard, so replicas add read capacity and only shards add write capacity. Everything else, including partitioned NoSQL stores, caches and queues, serves reads and writes on every replica.

    A node's request utilisation is the busier side for a single-primary store (replicas and primaries work in parallel) and the sum of both shares for everything else (they share the same machines). Its utilisation is the larger of that and its bandwidth utilisation (1.6). At or above 100% the node is saturated: every latency requirement of a use case that sends it load fails. Above 70% it is shown as hot.

    read capacity  = reads × replicas × shards
    write capacity = writes × shards              single-primary
                   = writes × replicas × shards   everything else
    
    ρ_read  = read load  ÷ read capacity
    ρ_write = write load ÷ write capacity
    
    ρ_req = max(ρ_read, ρ_write)   single-primary
          = ρ_read + ρ_write       everything else
    
    ρ = max(ρ_req, ρ_bandwidth)
  4. 1.4 Utilisation turns into queueing

    Each node has a base latency per call when idle (from its profile, below). Under load a request may wait for a free server. The model treats each replica of a shard as one server whose service time is that base latency, and uses the M/M/c queue: with c servers at utilisation ρ, a request waits with the Erlang C probability C, on average base ÷ (c(1 − ρ)). Load is spread evenly over shards and a key's requests can only go to its own shard, so c is one shard's replicas; writes that bind a single-primary store queue for its one primary.

    One server gives the classic 1 ÷ (1 − ρ) slowdown. A pool waits much less at the same utilisation: at 50%, one server waits half the time, four servers 17%. Utilisation is capped at 95% in this formula, so a saturated node has a finite latency (at most 20× its base) and fails its requirements through saturation instead.

    C      = Erlang C(c, ρ)              ρ capped at 0.95
    hop(n) = base + C × base ÷ (c (1 − ρ))
           = base ÷ (1 − ρ)                 for c = 1
    Slowdown over the idle latency by utilisation
    ρ10%50%70%90%≥ 95%
    1 server 1.1× 2× 3.3× 10× 20×
    4 servers 1× 1.09× 1.36× 2.97× 5.46×
  5. 1.5 Latency and percentiles

    A scenario's mean latency is the sum of the hops on the synchronous critical path of its entry request: requests sent before the entry request is answered, by the entry's sender or by a node still working on it synchronously. A par group counts its slowest member; an async send (->>) counts its own hop but not the receiver's work; a failed call (-x) costs a fixed timeout of 1000 ms instead of the target's latency, or what capacity { db timeout 200ms } sets for that node.

    Percentiles are computed, not measured. Each hop's time is a fixed part (half its service time) plus an exponential tail that carries the rest of its mean: the other half of the service time and all of the queueing. An idle hop's p99 is 2.8× its mean; one server at 50% makes it 3.7×, at 90% 4.4×. Timeouts and transfer time are fixed: they add once to every percentile. A path takes every hop at the same quantile, which errs on the slow side for long paths.

    A use case's percentile is the percentile of all its requests together: the scenarios mixed by their shares. With 10% cache misses, p99 is close to the miss path's p90; it moves smoothly as a share grows, without jumps. Both are solved numerically, the same way every time.

    item   = fixed + Exp(tail)
    fixed  = base ÷ 2 + transfer   (a -x: the timeout, no tail)
    tail   = hop(target) − base ÷ 2
    
    p_q(s) = Σ over the parts of the critical path of
               max over the items of the part of
                 fixed + tail × ln(1 ÷ (1 − q))
    
    p_q(U) = the t with Σ share(s) × P(s ≤ t) = q
    Percentiles of an idle hop over its mean
    idle hopp50p90p95p99p999
    ÷ mean 0.85× 1.65× 2× 2.8× 3.95×
  6. 1.6 Fan-out, payload size, bandwidth and egress

    A label that starts with x200 means the step happens 200 times per request: load counts 200 calls; latency counts the step once (as if the calls were batched or parallel), but its 200 payloads all cross the link.

    A label with ~2MB gives the step a payload: a read's payload travels back from the node it asks, a write's goes from the sender. It adds transfer time at the slower of the two ends' per-replica bandwidth (client 10 MB/s, edge, CDN, load balancer and gateway 1,000 MB/s, service 200 MB/s, everything else 100 MB/s). The bytes also fill the bandwidth of the nodes you run: a service whose payloads exceed its replicas' bandwidth is saturated. Object storage and CDNs scale out behind one name, so for them bandwidth is only the speed of one transfer.

    Payloads sent to a client or a third party leave your network and are charged as internet egress by the node that sends them: $0.09/GB from anything you run, $0.02/GB from a CDN. Traffic between your own nodes (storage to a service, a CDN filling from storage) is free, and what clients and third parties send is not your bill.

    transfer(step) = N × size ÷ min(bandwidth(from), bandwidth(to))
    
    ρ_bandwidth(n) = payload bytes in and out per second
                     ÷ (bandwidth × replicas × shards)
    
    egress GB/month = R × share × N × size × 2,592,000 s ÷ 10⁹
                      for payloads sent to a client or a third party
    egress cost     = egress GB × price of the node that sends it
  7. 1.7 Availability

    Each replica is up a fixed fraction of the time (profile below), independently of every other. A node is up when any of its replicas is. A use case is up when every node on the synchronous path of its main scenario (the one with the largest share) is up, except that a node with a fallback only takes the use case down when its fallback path is down too.

    A write to a single-primary store needs its primary. With a second replica, failover promotes one, but writes wait out the promotion: the model keeps 10% of the primary's downtime, so a PostgreSQL primary with replicas takes writes 99.995% of the time, alone 99.95%. Reads still count every replica.

    A fallback is a success scenario that calls the node with -x before the entry response and completes without a successful synchronous call to it. Load and saturation do not affect availability.

    A(n)       = 1 − (1 − a(n))^replicas
    A_write(n) = 1 − (1 − a(n)) × 0.1      single-primary, 2+ replicas
               = a(n)                      single-primary, 1 replica
    
    A(U) = Π over nodes n on the main path of
             A(n), or A_write(n) when the main path writes to it
             1 − (1 − A(n)) × (1 − A(fallback))     with a fallback
  8. 1.8 Cost

    Each replica of each shard costs a flat monthly price from its profile, whatever its load, plus the egress above, for a 30-day month. Clients, DNS and external systems cost nothing. There are no instance sizes, tiers, reservations or per-request prices.

    cost = Σ over nodes  replicas × shards × cost(n) + egress cost(n)
  9. 1.9 Failures: survive, fallbacks, single points of failure

    survive any node failure takes every node you run (not clients, external systems or DNS) and removes one instance. With two or more replicas, the whole design is analysed again with one replica fewer: the node must not saturate, and every latency requirement that held must still hold. On a sharded store the instance comes out of one shard; keys cannot move to another shard, so that shard's numbers are what count. A single-primary store that loses its primary promotes a replica: writes keep a primary and reads lose one replica, the same as losing a replica. With one replica, every use case whose success scenarios need the node must have a fallback scenario for it. survive failure of <selector> does the same for the selected nodes.

    The Analysis tab lists as single points of failure the single-replica nodes that some use case needs and has no fallback for. A failed call (-x) in a scenario is how a design says "and this is what happens when that node is down".

    x2+
    with one replica fewer, the node does not saturate and the latency limits still hold
    x1
    every use case that needs it has a fallback scenario
    Skipped
    clients, external systems, DNS, notes
  10. 1.10 Tech stacks, kinds and consistency

    A node's tech stack picks its kind and its profile. The catalog knows 205 techs with the names people write as aliases ([S3], [Postgres], [ALB], [k8s]), ignoring case, punctuation and a trailing version. A tech it does not know is a warning, and the node is simulated as the kind its name suggests ([TigerBeetle DB] is a database), else the kind of the closest catalog tech, else a service: never a free, infinitely fast client.

    The edge splits into cdn, loadbalancer, gateway and dns, each with its own numbers; a firewall ([WAF]) is a plain edge, and any edge selects all of them. DNS is off the request path: it adds no latency, never fails and costs nothing. Every data store has a consistency, strong or eventual, used only by the selectors any strong store and any eventual store in tests; replication lag and staleness are not simulated.

    [Postgress]         did you mean PostgreSQL? → database
    [TigerBeetle DB]    unknown → database (from its name)
    [Cobol Mainframe]   unknown → service
    
    any edge            edge, cdn, loadbalancer, gateway, dns
    any strong store    relational, MongoDB, Spanner, queues …
    any eventual store  caches, DynamoDB, Cassandra, search, storage …

Default profiles, per replica

Teaching values, right to an order of magnitude, not benchmarks. capacity { … } overrides any of them per node (see How to read the results).

Kind Examples Reads/s Writes/s Latency Avail. Cost/mo Bandwidth Egress Durable Consistency
clientActor, shapes without a tech∞∞0 ms100%$010 MB/s—no—
cdnCloudFront, Azure CDN, Cloud CDN, Front Door200k200k5 ms99.99%$1001,000 MB/s$0.02/GBno—
edgeWAF, AWS WAF, Global Accelerator100k100k2 ms99.99%$501,000 MB/s$0.09/GBno—
loadbalancerAWS Load Balancer (ALB, NLB), nginx, Envoy, HAProxy100k100k2 ms99.99%$501,000 MB/s$0.09/GBno—
gatewayAWS API Gateway, Azure API Management, Kong10k10k10 ms99.95%$1001,000 MB/s$0.09/GBno—
dnsRoute53, Azure DNS, Cloud DNS∞∞0 ms100%$01,000 MB/s—no—
serviceService, REST API, gRPC, Spring Boot, Go, Node.js, ECS, EKS, Kubernetes2k2k10 ms99.5%$100200 MB/s$0.09/GBno—
functionLambda, Cloud Functions, Cloud Run10k10k25 ms99.95%$200100 MB/s$0.09/GBno—
cacheRedis, Valkey, ElastiCache, Memcached100k100k1 ms99.9%$150100 MB/s$0.09/GBnoeventual
database, relationalPostgreSQL, MySQL, Aurora, RDS, SQL Server, Oracle20k5k5 ms99.95%$400100 MB/s$0.09/GByesstrong
database, NoSQLDynamoDB, Cassandra, ScyllaDB, Cosmos DB20k20k5 ms99.99%$500100 MB/s$0.09/GByeseventual
database, NoSQLMongoDB, CockroachDB, Bigtable, Spanner20k20k5 ms99.99%$500100 MB/s$0.09/GByesstrong
searchElasticsearch, OpenSearch, Solr3k3k15 ms99.9%$400100 MB/s$0.09/GByeseventual
analyticsBigQuery, Snowflake, Redshift, ClickHouse200200500 ms99.9%$300100 MB/s$0.09/GByeseventual
queueKafka, Kinesis, SQS, SNS, Pub/Sub, RabbitMQ50k50k5 ms99.99%$200100 MB/s$0.09/GByesstrong
storageS3, GCS, Azure Blob, MinIO5k5k30 ms99.99%$50100 MB/s$0.09/GByeseventual
externalStripe, Twilio, SendGrid, third-party APIs1k1k200 ms99.9%$0100 MB/s—no—

Relational databases take 5k writes per second per shard (on the primary), not per replica. Egress is the price of data a node sends to clients and third parties. A tech stack the catalog does not know gets the row of the kind its name suggests (a service when nothing does) and a warning that names the closest known tech; to keep a product name the catalog lacks without the warning, use the generic tech of its kind ([Database], [Message Queue]) and put the product in the node's name.

02

Worked examples

Every number in this section comes from running the simulation on the source shown next to it, and a test fails if the code and the page disagree. Paste a source into the editor to see the same numbers in its Analysis and Tests tabs.

2.1 A cached read path

title "Feed"

user  "User"       [Actor]
api   "Feed API"   [REST API]   x6
cache "Feed cache" [Redis]      x2
db    "Feed DB"    [PostgreSQL] x2

user -> api  : HTTPS
api -> cache : RESP
api -> db    : SQL

usecase "Read feed" {
  user -> api : GET /feed
  alt "Cache hit" {
    api  -> cache : GET feed:{user}
    api --> user  : 200
  } alt "Cache miss" {
    api  -> cache : GET feed:{user}
    api  -> db    : SELECT posts
    api  -> cache : SET feed:{user}
    api --> user  : 200
  }
}

usecase "Post" {
  user -> api  : POST /posts
  api  -> db   : INSERT post
  api --> user : 201
}

traffic {
  "Read feed" 6k rps mix "Cache hit" 90%, "Cache miss" 10%
  "Post"      1k rps
}

requirements {
  p99 "Read feed" < 60ms
  availability "Read feed" >= 99.9%
  durable "Post"
  survive any node failure
  cost <= 2500 usd/month
}

Load. api receives every entry request: 6k reads and 1k writes a second against 12k rps (6 × 2k), so ρ = 50% + 8% = 58%. db gets only the misses (10% of 6k = 600 rps SELECTs) and every post (1k rps INSERTs). Its two replicas serve 40k rps of reads but its one primary takes 5k rps of writes, so writes set its utilisation: 20%.

Latency. The six API replicas at 58% queue as one pool of 6 servers: a request waits only 18% of the time, so the API's 10 ms becomes 10.7 ms (one server at 58% would take 24 ms). The cache stays at 1 ms and the database at 6.3 ms; its writes queue for the one primary. The hit path is api + cache, mean 11.7 ms; the miss path adds the SELECT and the SET, mean 19 ms.

Percentiles. p99 is taken over all reads, 90% hits and 10% misses mixed: 38.6 ms, above the hit path's own p99 (34.1 ms) and below the miss path's (56.7 ms). p50 is 10.3 ms, a little under the hit path's mean: half of each service time is fixed, and the median of the exponential rest is below its mean.

Cost. 6 × $100 + 2 × $150 + 2 × $400 = $1,700 a month.

pass
p99 of Read feed is 38.6 ms (limit 60 ms)
pass
availability of Read feed is 99.9999% (limit 99.9%)
pass
"Post" writes to a durable store before responding
pass
Every use case keeps working after losing any of 3 nodes
pass
Total cost is $1,700/month (limit $2,500/month)

2.2 Replicas do not scale writes

title "Orders"

client "Client"    [Actor]
api    "Order API" [REST API]   x8
db     "Orders DB" [PostgreSQL] x3

client -> api : HTTPS
api    -> db  : SQL

usecase "Place order" {
  client -> api    : POST /orders
  api    -> db     : INSERT order
  api   --> client : 201
}

traffic {
  "Place order" 6k rps
}

requirements {
  p99 "Place order" < 100ms
}

Then add two shards:

capacity {
  db shards 2
}

Saturated. 6k INSERTs a second reach a PostgreSQL primary that takes 5k. Three replicas give 60k rps of read capacity, none of it usable for writes, so the database runs at 120% and the requirement fails however much time is left:

fail
p99 of Place order: db is saturated (reads 0 rps of 60k rps, 0%; writes 6k rps of 5k rps, 120%) (limit 100 ms)Writes to db (PostgreSQL) need 6k rps and each primary takes 5k rps; read replicas do not add write capacity. Add shards (capacity { db shards 2 } keeps writes under 70%), move the data to a partitioned store (DynamoDB, Cassandra), or batch writes through a queue

Sharded. With shards 2 there are two primaries, each with three replicas: write capacity 10k rps, utilisation 60%, and the database hop takes 12.5 ms. The cost doubles too: 3 replicas × 2 shards × $400 = $2,400.

pass
p99 of Place order is 76.7 ms (limit 100 ms)

Moving the table to a partitioned store such as DynamoDB, which scales writes with replicas, would pass as well. In a practice problem shards is the one capacity setting you may write yourself, because how many shards to run is a design decision.

2.3 Bytes, bandwidth and egress

title "Photos through the API"

viewer "Viewer"    [Actor]
api    "Photo API" [REST API] x2
blobs  "Photos"    [AWS S3]   x2

usecase "View photo" {
  viewer -> api    : ~2MB GET /photos/{id}
  api    -> blobs  : ~2MB GET photo
  api   --> viewer : 200
}

traffic {
  "View photo" 20 rps
}
title "Photos through a CDN"

viewer "Viewer" [Actor]
cdn    "CDN"    [AWS CloudFront] x2
blobs  "Photos" [AWS S3]         x2

usecase "View photo" {
  viewer -> cdn : ~2MB GET /photos/{id}
  alt "Edge hit" {
    cdn --> viewer : 200
  } alt "Edge miss" {
    cdn   -> blobs  : ~2MB GET photo
    cdn  --> viewer : 200
  }
}

traffic {
  "View photo" 20 rps mix "Edge hit" 95%, "Edge miss" 5%
}

Transfer time. A 2 MB answer to the viewer moves at the viewer's 10 MB/s: 200 ms, more than every node's own latency together, and a fixed cost that adds once to every percentile. Through the API the photo also crosses api ↔ storage at storage's 100 MB/s (20 ms): mean 260 ms, p50 254 ms and p99 334 ms. Through the CDN, hits take 205 ms and misses 255 ms; with 5% misses, p99 is 266 ms. Latency barely improves, because the viewer's own bandwidth dominates both designs.

Egress. 20 photos of 2 MB a second is 103,680 GB a month sent to viewers. The API sends them at the internet rate: $9,331 of egress and $9,631 in total; what storage sends the API stays inside and costs $0. Through the CDN, the edge sends the same bytes for $2,074, and filling it from storage is free: $2,374 in total.

03

Simplifications that may surprise you

These follow directly from the formulas above. None of them changes which of two designs is better in ordinary cases, but each one moves absolute numbers in a known direction.

Simplifications in the model and their effect
Slow hops come together pessimistic A path's percentile takes every hop at the same quantile, as if the slow requests of every node hit the same request. Real hops vary independently, so long paths have narrower tails than the model's. Kept because it is simple, deterministic and errs on the safe side; the sum of independent hops has no closed form.
A replica is one server either way Queueing treats each replica as one server whose service time is its base latency. A real replica serves many requests at once, so its latency stays flat longer and then climbs faster. The shape (more replicas queue less, one busy server queues a lot) is right; the exact curve is a teaching value.
Fan-out latency counted once optimistic x200 multiplies load and transfer, but adds one hop of latency. The slowest of 200 parallel calls is far slower than a typical one (tail at scale). Kept, because the same label also means a batched call, which really is one round trip.
Independent replicas optimistic 1 − (1 − a)^n assumes replicas never fail together: six service replicas at 99.5% each show as 99.999999%. A bad deploy, a zone outage or a shared dependency takes them all at once.
Failover is a fixed share either way Writes to a single-primary store with replicas lose 10% of the primary's downtime to failover, whatever the database or its settings. Real failovers take from seconds to minutes, and some lose the last writes.
Saturation does not spread optimistic A saturated node still passes its full load downstream and its latency stops growing at 20×. Real queues grow without bound, time out and push retries upstream. Kept on purpose: a design with a saturated node already fails, and throttling the load would hide the next bottleneck behind the first, so you would fix them one at a time.
Managed bandwidth is unlimited optimistic Object storage and CDNs never saturate on bytes; their bandwidth only sets the speed of one transfer. That is close to how they behave, but per-prefix request limits and origin shields are not modelled.

04

What is not modelled

Most omissions make the model optimistic: real systems are slower, less available and dearer than the numbers say. Treat a design that only just passes as one that would fail.

What the model leaves out and what that does to the numbers
Cold starts optimistic A function always takes its 25 ms. Real cold starts add hundreds of milliseconds to the tail, more for big runtimes.
GC pauses, CPU steal, noisy neighbours optimistic Every replica is equally fast all the time. Pauses and shared hardware fatten p99 and p999 well beyond the tails the model computes from queueing alone.
Network latency and jitter optimistic A hop costs the target's latency only; there is no separate round trip, no cross-zone or cross-region distance and no packet loss. Multi-region designs look as fast as single-zone ones.
Retries and thundering herds optimistic A failed call costs one timeout and the scenario moves on. Retries that double the load on a sick node, retry storms after an outage and a cold cache stampeding the database do not exist.
Partial and grey failures optimistic Nodes are up or down. A slow replica, a half-broken network path or a dependency that answers with errors at 2% is not representable, except as a scenario you write.
Cache hit rate either way The hit rate is whatever mix says. Warm-up, eviction, a cache emptied by a restart, and hit rates that change with traffic or data size are not modelled. Pick the rate you would actually measure, and model the cold case as its own scenario.
Hot keys and skew optimistic Load is spread evenly over every replica and shard. One celebrity account or one hot partition saturates a real shard long before the average does.
Traffic shape and autoscaling optimistic Traffic is one constant rate and replica counts are fixed. Peaks, daily cycles, bursts and the minutes an autoscaler needs to catch up are not modelled. Enter your peak rate, not your average.
Contention and per-request work either way Every request to a node costs the same capacity: a key lookup and a five-table join, a 1 KB and a 1 MB write. Locks, connection pools, transactions and replication work on followers are not modelled.
Queues and consumers either way Backlog, consumer lag, ordering, redelivery and dead letters are out of scope. A queue is a node that takes writes; its consumers are steps with load.
Time and consistency either way TTLs, expiry, replication lag and stale reads are not simulated; strong and eventual are labels for tests. Model expiry or a stale read as a scenario.
Third-party limits optimistic External systems are free, take 1k requests a second per replica and never rate-limit you.
Real cloud pricing either way One flat monthly price per replica. No instance sizes, reserved or spot pricing, per-request pricing (Lambda, DynamoDB on demand, API Gateway, S3 requests), storage at rest, free tiers or egress tiers. Traffic between your own nodes is free, also across zones, where clouds charge about $0.01/GB each way (optimistic, and small next to internet egress).

05

How to read the results

  1. 5.1 Compare designs, find the bottleneck

    The model is deterministic and its inputs are rough, so read its numbers relatively. Two designs under the same traffic and profiles compare fairly: a cache that takes p99 from 300 ms to 40 ms, a database that goes from 120% to 60%, a CDN that cuts egress by four. A 10% difference between two designs is inside the error of the teaching values.

    The utilisation column shows where the design gives out first. Raise the traffic until something turns amber (above 70%) or red (saturated): that node, and the lever the hint names (replicas, shards, a cache, async work), is what the design is really about.

    Do not size a fleet from these numbers or quote them as SLOs. A capacity plan needs load tests of your own services, your peak traffic and your cloud bill.

    Good for
    comparing designs, finding the first bottleneck, spotting single points of failure, sanity-checking a budget
    Not for
    capacity plans, SLO commitments, cost forecasts, tail-latency estimates
    Fixed
    no randomness, no warm-up: the same text always gives the same numbers
  2. 5.2 Calibrate with your own measurements

    capacity { … } replaces any default for one node, per replica. Once you have real numbers, put them in and the rest of the model works from them:

    • Rate (3k rps, or reads … writes …): the throughput at which one replica stops keeping up in a load test, since the model adds queueing delay on top. Not the rate at which latency is still fine.
    • Latency: the median per call at low load, measured from the caller, so it includes the network.
    • Availability: per replica, from your provider's SLA or your incident history.
    • Cost: what one replica costs on your bill. Egress and bandwidth: your prices and your instance's network.
    • Shards, durable, consistency: how the store is actually set up.

    Calibrated rates fix the capacity side; the omissions in section 04 still apply, so keep headroom.

    capacity {
      api   3k rps latency 12ms cost 140 usd/month
      db    reads 25k rps writes 6k rps availability 99.95%
      users shards 4 consistency strong
      blobs bandwidth 500 MB/s egress 0.05 usd/GB
    }

06

Practice verdicts

  1. 6.1 What "Solved" means

    A practice problem is solved when every requirement and every test in the problem's given file passes against your design, with the same simulation as the editor. The problem fixes the traffic, the limits and any capacity numbers; a capacity line in your own file is an error, except shards <n>, so you cannot buy your way out with a bigger node.

    Passing means the design is sound under this model: the requests flow the right way, nothing saturates at the given traffic with the default numbers, every node you run can fail without breaking a use case, and the bill fits the budget. It does not mean the system would hold up in production; section 04 is still true.

    Latency
    p99 limits fail when the path is too long, a node queues, or any node a use case loads is saturated
    Failure
    survive and availability fail on single points of failure
    Cost
    the budget fails designs that throw replicas at the problem
    Flow
    tests check the key idea: cache before database, write before responding, never wait for a queue
  2. 6.2 Why brute force fails

    Each problem's limits are set between the reference solution and the shortcuts a candidate might try: loose enough that the intended design passes with room, tight enough that the shortcut does not. More replicas fix saturation but break the cost budget; skipping the cache saves its cost but loads the database and the p99; a load balancer in front of storage fails the flow test that asks for a CDN.

    Every problem keeps those shortcuts as wrong/*.proschi designs, each naming the tests it must fail. proschi problem check runs them in CI, so a change to the model or the limits that lets a shortcut pass is caught before it ships. How problems are calibrated.

    # expect-fail: p99 of Hold seat < 130 ms
    # Scales the relational database with read replicas instead of shards.
    import "problem.proschi"
    
    # … the rest of the design …
    db "Seats DB" [PostgreSQL] x4

View the source on GitHub