Review cards · 18 cards

Observability

Metrics, logs and traces, SLOs and error budgets, percentiles and alerting on what users feel.

Train observability in daily review

Cards

  1. An endpoint's average latency is a steady 50 ms. Why might users still complain that it is slow? easy Multiple choice
  2. Metrics, logs and traces: what question does each answer best? easy Flashcard
  3. For every service, the RED method tracks the _____ of requests, the _____ and their _____. For resources like CPUs, disks and pools, the USE method tracks utilization, saturation and errors. easy Fill in the blank
  4. How do an SLI, an SLO and an SLA relate? easy Flashcard
  5. Why write logs as structured events (JSON fields such as user_id, trace_id, duration_ms) rather than free-text lines? easy Flashcard
  6. Three servers behind a load balancer report p99 latencies of 100 ms, 200 ms and 900 ms. What is the p99 for the whole service? medium Multiple choice
  7. Service A calls B over HTTP and also sends it jobs through a queue, but B's spans show up as separate traces instead of under A's. What is the likely cause? medium Multiple choice
  8. A team measures its 99.99% SLO over a rolling 7-day window. How much total outage fits in one window? medium Estimate
  9. Why record latency as a histogram (counts per bucket) rather than as a p99 each server computes for itself? medium Flashcard
  10. Which label on a http_requests_total counter is most likely to overload the metrics database? medium Multiple choice
  11. Why do teams avoid setting an SLO of 100%? medium Multiple choice
  12. Pages are slow only for users on mobile networks in one country. Which monitoring is most likely to catch it? medium Multiple choice
  13. An SLO of 99.95% availability over a 30-day window: how many minutes of total outage does its error budget hold? medium Estimate
  14. All the spans of one request share a _____. Each span has its own _____ and records its _____'s, which is how the tracing system rebuilds the call tree. medium Fill in the blank
  15. Why alert on the error budget's burn rate over two windows rather than on "error rate above 1% for 5 minutes"? hard Flashcard
  16. A service has a 99.9% SLO over 30 days. A bad deploy makes 1.44% of requests fail and stays out. If nothing changes, how long until the whole month's error budget is gone? hard Estimate
  17. A page calls 100 backend servers in parallel and waits for all of them. Each server answers in over 1 second for 1% of its requests (its p99 is 1 s). What share of page loads take over 1 second? hard Estimate
  18. Head-based or tail-based trace sampling: what does each trade? hard Flashcard

More topics