Review cards · 14 cards
Availability
Nines, redundancy, failover and how availability combines across components.
Cards
- A service promises 99.99% availability. How much downtime does that allow in a year?
- How much recent data you can afford to lose in a disaster, measured in time, is the _____. How long you can afford to be down is the _____.
- A service promises 99.9% availability. How much downtime does that allow in a 30-day month?
- Active-active versus active-passive across two sites: what does each cost?
- A request passes through two components in a row, each 99.9% available. What is the availability of the whole path?
- Your service calls a provider that is only 99% available. How can your service be more available than it?
- The math said two replicas give about 99.9999%, but both went down together. Which assumption broke?
- A service has a 99.95% success-rate SLO over 30 days and serves 1 million requests a day. How many failed requests can it afford in that period?
- Many outages start with a change. How do you limit the damage of a bad deploy?
- Peak load needs 6 servers. They are spread evenly over 3 availability zones, and the service must survive losing a whole zone. How many servers do you run?
- Which setup loses no acknowledged writes if the database's availability zone is destroyed?
- What is a cell-based architecture, and what does it buy?
- Should a load balancer's health check also check the database? What can go wrong?
- Three independent replicas are each 95% available, and the service is up while any one of them is. How many minutes a month (30 days) is it down?
More topics
- Estimation 21 cards
- Networking 16 cards
- API design 17 cards
- Caching 21 cards
- Databases 22 cards
- Replication 15 cards
- Sharding 18 cards
- Consistency 19 cards
- Queues 18 cards
- Streaming 18 cards
- Resilience 16 cards
- Storage 14 cards
- Realtime 15 cards
- Data structures 16 cards
- Security 17 cards
- Observability 18 cards
- Coordination 16 cards