System design practice

File Storage

medium storage cdn egress async

Pre-signed uploads: the API never touches the bytes.

Solve it in your browser Read the lesson first

Design a small Dropbox: users upload files, list their folders and download files again, from any device. Files are anything from a 10 KB note to a 2 GB video, so their bytes must never stream through your own servers: the API hands out pre-signed URLs and the client talks to object storage (for uploads) and to a CDN (for downloads) directly.

Functional requirements

Use these use case and scenario names exactly: the traffic, requirements and tests in problem.proschi refer to them.

Scale

Constraints

What is given

problem.proschi declares the user and the bucket blobs (S3, two partitions) and holds the traffic, requirements and tests. Add the API, the metadata store, the CDN, how a stored object marks its file ready, and the four use cases.

How your design is checked

You write the design as text in Proschi. Tests run in your browser: a simulation of the traffic above checks latency, availability, cost and what happens when a machine fails. How the simulation works.

Start designing

Lesson · 13 min read

Learn it: File Storage #

File Storage: keep the bytes off your servers #

A small Dropbox sounds like a CRUD app (create, read, update, delete) with big rows. It is not. The moment files range from a 10 KB note to a 2 GB video, the design question changes from "where do I store this?" to "which machines should the bytes travel through, and who pays for them on the way out?" This lesson builds the answer: pre-signed URLs for uploads, bucket events to finish them, and a CDN for downloads.

What you'll learn #

The problem, explained #

Who uses it. People with several devices who want their files everywhere: upload a report on a laptop, open it on a phone.

Functional requirements.

Non-functional requirements. p99 under 200 ms for listing, starting an upload and downloading, and under 220 ms for the upload itself, transfer time included. Every use case available 99.9%. Uploads are acknowledged only when the bytes are durable; the pending record exists before the URL is handed out. Clients never reach the metadata store. Any single machine can fail. At most $45,000 a month, egress included.

What is given. given.proschi declares the user and the bucket blobs (S3 with two partitions). Object storage is given because it is not the interesting choice: every cloud has one, and it scales by itself. The interesting choices are everything around it.

What the tests check.

Back-of-the-envelope #

The problem gives the rates and an average file of 1 MB. A month in the model is 30 days, 2,592,000 seconds.

QuantityArithmeticResult
Download bandwidth500 rps × 1 MB500 MB/s
Downloaded per month500 MB/s × 2,592,000 s1,296,000 GB ≈ 1.3 PB
Egress if the bucket sends it1,296,000 GB × $0.09about $116,600/month
Egress if the CDN sends it1,296,000 GB × $0.02about $25,900/month
Bucket reads with 90% CDN hits10% × 500 rps50 rps
Upload bandwidth (ingress, free)200 rps × 1 MB200 MB/s
New data per month200 MB/s × 2,592,000 sabout 518 TB
API requests2k list + 200 start2.2k rps
Metadata writes200 inserts + 200 updates400 rps
Transfer time for 1 MB at the user's 10 MB/s1 MB ÷ 10 MB/s100 ms

Three things jump out.

Egress decides the design. Sending 1.3 PB from the bucket alone costs more than twice the whole budget. Sending it from a CDN fits with room. No amount of clever compute changes that, so the CDN is not an optimisation here; it is a requirement hiding in the cost line.

Transfer time sets the latency floor. A 1 MB file over a 10 MB/s connection takes 100 ms no matter what. In Proschi, payload transfer is a fixed cost that adds once to every percentile, so a 200 ms p99 budget has only about 100 ms left for everything else. Every extra hop that also carries the bytes eats into it.

Storage grows fast. About half a petabyte of new files a month. Proschi's cost model has no storage-at-rest price, but in an interview you should mention lifecycle rules (move cold files to cheaper storage classes) and deduplication.

How the model sees each piece. A load balancer takes 100k rps per replica, a service 2k, PostgreSQL 20k reads but 5k writes on its single primary, a queue 50k, a function 10k, storage 5k. With 2.2k API requests a second, about two service replicas would be fully busy. Divide by your target utilisation, and check that one replica fewer still survives. Metadata writes are far below one primary's capacity, so no sharding is needed. Object storage and CDNs "scale out behind one name": the model never saturates them on bandwidth, which matches how they behave in reality.

Availability. A write to a single-primary database with two replicas is modelled at 99.995% (failover costs 10% of the primary's downtime). That is the weakest link for Start upload, still well above 99.9%.

Concepts #

Object storage and pre-signed URLs #

Object storage (S3, GCS, Azure Blob) stores immutable blobs under keys. It is cheap per byte, extremely durable, and scales request rates and bandwidth for you. What it is not: a database you can query, or a file system with cheap renames.

A pre-signed URL is a URL that carries a signature made with your credentials. It allows one operation (say, a PUT to one key) and expires after a few minutes. Your API decides who may upload what, signs the URL locally (no network call), and hands it to the client. The client then talks to the bucket directly. Your servers never see the bytes, so they need neither the bandwidth nor the long-lived connections a 2 GB upload would hold.

Trade-offs: you cannot inspect the bytes on the way in, so virus scanning and transcoding happen later, on an event. URLs can leak until they expire, so keep expiry short and scope each URL to one key. And something must tell your system that the upload happened, which is the next concept.

When not to use it: tiny payloads that are really part of an API call (an avatar crop, a 2 KB JSON), where an extra round trip costs more than proxying.

title "Direct-to-bucket upload"

phone  "Phone"     [Actor]
api    "Media API" [REST API]   x2
db     "Media DB"  [PostgreSQL] x2
bucket "Bucket"    [AWS S3]     x2

phone -> api    : HTTPS
phone -> bucket : PUT pre-signed
api   -> db     : SQL

usecase "Get upload URL" {
  phone -> api   : POST /media
  api   -> db    : INSERT media status=pending
  db   --> api   : ok
  api  --> phone : 201 {"putUrl": "https://bucket.example.com/m1?sig=..."}
}

usecase "Put bytes" {
  phone   -> bucket : ~5MB PUT /m1?sig=...
  bucket --> phone  : 200
}

Event-driven completion #

After the PUT, someone must flip the record from pending to ready. The naive answer is "the client calls /complete". But phones lose signal, laptops close, apps crash. A client that disappears leaves a file pending forever.

The robust answer: the bucket itself announces new objects. S3 event notifications can deliver an ObjectCreated event to SQS, SNS, Lambda or EventBridge. Put a queue in between so events survive a slow or failing consumer, and let a worker or function mark the file ready. If the worker crashes, the queue redelivers. Because delivery is at-least-once, the update must be idempotent: setting status=ready twice is harmless, which is why it is a good shape for this step.

Trade-offs: a short delay between the PUT finishing and the file appearing as ready; the client should poll or be pushed the state change. When not to use it: if you need the upload's result synchronously (for example, a server-side validation that must reject the file before the user moves on), you need a synchronous step, usually a proxy or a post-upload check the client waits for.

title "Object events"

bucket "Bucket"      [AWS S3]     x2
events "Object Feed" [AWS SQS]    x2
fn     "Indexer"     [AWS Lambda] x2
db     "Catalog"     [DynamoDB]   x2

bucket -> events : ObjectCreated
events -> fn     : trigger
fn     -> db     : write

usecase "Index new object" {
  bucket ->> events : ObjectCreated img/42.jpg
  events ->> fn     : ObjectCreated img/42.jpg
  fn      -> db     : UPDATE item 42 status=indexed
  db     --> fn     : ok
}

CDNs for large downloads #

A CDN is a fleet of caching proxies near users. On a hit, the edge serves the file from its cache; on a miss, it fetches from the origin (here, the bucket), keeps a copy and serves it. This is a pull CDN, the kind the System Design Primer recommends for heavy traffic.

Two different benefits, often confused:

When not to use a CDN: highly personalised content that is never reused, or tiny traffic where the fixed cost outweighs savings. Private files are fine on a CDN as long as links are signed and short-lived.

Designing it step by step #

1. Scope. Ask: file sizes (up to 2 GB, so streaming through servers is out), read/write ratio (downloads plus listings far outnumber uploads), whether files are shared (yes, by link, hence signed URLs), and the cost constraint. Confirm sync, versioning and conflict resolution are out of scope; they are great follow-ups but not this problem.

2. High-level design. Separate metadata from content:

Then sketch the four use cases. List files: user → load balancer → API → database. Start upload: the same path, with an insert of a pending row and a 201. Upload: user → bucket, then the bucket's event travels to a finisher. Download: user → CDN, and on a miss CDN → bucket.

3. Deep dive.

4. Wrap-up. Walk the requirements: upload p99 is the 100 ms transfer plus the bucket's write; download p99 is the transfer plus a CDN hop, with 10% misses adding the origin fetch; nothing is a single replica; the client never sees the database. Then list improvements: multipart uploads for large files (resume after a dropped connection), content hashing for deduplication, a sweeper that deletes pending rows whose upload never arrived, and lifecycle tiers for cold data.

Common mistakes #

Proxying the upload through the API (wrong/upload-through-api). The most common instinct: the client sends bytes to your API, which forwards them to the bucket. In the real world, each 2 GB upload pins an API connection for minutes and your API fleet scales with bandwidth instead of requests. In the model, the extra hops (load balancer, API, and a second transfer between API and bucket) push the upload's p99 over 220 ms. It fails File bytes never pass through your servers and p99 of Upload < 220 ms.

The client marks its own upload complete (wrong/client-marks-upload-complete). After the PUT, the client calls POST /files/{id}/complete. It works in the demo, then leaves orphaned pending files every time a phone loses signal at the wrong moment. It fails Uploads are finished without the client, because the user calls the load balancer during Upload.

Downloading straight from the bucket (wrong/download-from-bucket). Signed bucket URLs are simple and correct, but every byte leaves at $0.09/GB. The model puts the bill around $119k a month, far over $45k. It fails Downloads are served by the CDN and cost ≤ $45,000/month.

A load balancer where the CDN should be (wrong/load-balancer-instead-of-cdn). A load balancer distributes traffic but caches nothing, so every download still reaches the origin and egress is still billed at the higher rate. It fails Downloads are served by the CDN, because the selector any cdn does not match a load balancer.

Other classic mistakes.

In the interview #

Open with the two numbers that drive everything: "Downloads move about 1.3 PB a month, and a 1 MB file takes 100 ms on a user's connection." Then state the principle: "Our servers handle metadata and authorisation; bytes go client ↔ bucket ↔ CDN." Draw the metadata path and the byte path in different colours if you can.

Likely follow-ups:

Further reading #

Now design it

More system design problems