Review cards · 18 cards
Observability
Metrics, logs and traces, SLOs and error budgets, percentiles and alerting on what users feel.
Cards
- An endpoint's average latency is a steady 50 ms. Why might users still complain that it is slow?
- Metrics, logs and traces: what question does each answer best?
- For every service, the RED method tracks the _____ of requests, the _____ and their _____. For resources like CPUs, disks and pools, the USE method tracks utilization, saturation and errors.
- How do an SLI, an SLO and an SLA relate?
- Why write logs as structured events (JSON fields such as user_id, trace_id, duration_ms) rather than free-text lines?
- Three servers behind a load balancer report p99 latencies of 100 ms, 200 ms and 900 ms. What is the p99 for the whole service?
- Service A calls B over HTTP and also sends it jobs through a queue, but B's spans show up as separate traces instead of under A's. What is the likely cause?
- A team measures its 99.99% SLO over a rolling 7-day window. How much total outage fits in one window?
- Why record latency as a histogram (counts per bucket) rather than as a p99 each server computes for itself?
- Which label on a http_requests_total counter is most likely to overload the metrics database?
- Why do teams avoid setting an SLO of 100%?
- Pages are slow only for users on mobile networks in one country. Which monitoring is most likely to catch it?
- An SLO of 99.95% availability over a 30-day window: how many minutes of total outage does its error budget hold?
- All the spans of one request share a _____. Each span has its own _____ and records its _____'s, which is how the tracing system rebuilds the call tree.
- Why alert on the error budget's burn rate over two windows rather than on "error rate above 1% for 5 minutes"?
- A service has a 99.9% SLO over 30 days. A bad deploy makes 1.44% of requests fail and stays out. If nothing changes, how long until the whole month's error budget is gone?
- A page calls 100 backend servers in parallel and waits for all of them. Each server answers in over 1 second for 1% of its requests (its p99 is 1 s). What share of page loads take over 1 second?
- Head-based or tail-based trace sampling: what does each trade?
More topics
- Estimation 21 cards
- Networking 16 cards
- API design 17 cards
- Caching 21 cards
- Databases 22 cards
- Replication 15 cards
- Sharding 18 cards
- Consistency 19 cards
- Queues 18 cards
- Streaming 18 cards
- Availability 14 cards
- Resilience 16 cards
- Storage 14 cards
- Realtime 15 cards
- Data structures 16 cards
- Security 17 cards
- Coordination 16 cards