Platform and scale

SLOs people actually use

Three signals per service, one dashboard, and alerts that correspond to someone being paged.

Observability 9 min read

Most Kubernetes monitoring setups collect everything and tell you nothing. The failure is not technical — it is that nobody decided what "working" means before building the dashboards.

Delete half your alerts first

Before adding anything, go through what fires today and ask one question of each: when this fired, did a human do something? If the honest answer is no, it is not an alert. It is a dashboard panel at best.

An alert nobody acts on is worse than no alert, because it trains the team to ignore the channel where the real one will arrive. The strongest predictor of whether monitoring works is not coverage, it is how many pages people trust.

Three signals is enough

  • Availability — the fraction of requests that did not fail.
  • Latency — the fraction served faster than a threshold you chose deliberately.
  • Saturation — how close the thing is to a limit, so you have warning before the first two move.
# availability, as a ratio - not a count of errors
sum(rate(http_requests_total{status!~"5.."}[5m]))
  / sum(rate(http_requests_total[5m]))

# latency, as a ratio of requests inside the target
sum(rate(http_request_duration_seconds_bucket{le="0.5"}[5m]))
  / sum(rate(http_request_duration_seconds_count[5m]))
Express latency as "what fraction was fast enough" rather than as a percentile. Histogram-derived percentiles cannot be averaged or aggregated correctly across instances, and a ratio can. It also matches how you state an objective: 99% of requests under 500 ms.

Alert on burn rate, not on breach

A 99.9% monthly objective gives you about 43 minutes of error budget. Alerting the moment you dip below 99.9% in a five-minute window is pure noise. Alerting when you are consuming the budget fast enough to exhaust it is the signal.

# fast burn: 2% of a 30-day budget in an hour -> page
- alert: ErrorBudgetBurningFast
  expr: |
    (1 - (sum(rate(http_requests_total{status!~"5.."}[1h]))
          / sum(rate(http_requests_total[1h])))) > 14.4 * 0.001
  for: 2m
  labels: { severity: page }

# slow burn: on course to exhaust it this week -> ticket, not a page
- alert: ErrorBudgetBurningSlow
  expr: |
    (1 - (sum(rate(http_requests_total{status!~"5.."}[6h]))
          / sum(rate(http_requests_total[6h])))) > 6 * 0.001
  for: 30m
  labels: { severity: ticket }

Two severities, two response paths. The fast burn wakes someone; the slow burn becomes work on Monday. That distinction is what makes an on-call rotation survivable.

Cluster-level alerts worth keeping

  • A node NotReady for more than five minutes.
  • Pods in CrashLoopBackOff for more than fifteen.
  • PersistentVolume above 85% full — with hours of warning, not minutes.
  • Certificates expiring within fourteen days.
  • A Deployment with fewer ready replicas than its PDB requires.

Notably absent: anything about individual pod restarts, CPU above a threshold, or memory above a threshold. Those are symptoms, they fire constantly, and they are what dashboards are for.

Cardinality is how this gets expensive

One label with unbounded values will cost more than the rest of your monitoring combined. A user_id, a request path with an id in it, or a pod name on a high-churn Deployment creates a new time series per value, forever. Normalise paths to route templates before they reach the metric, and keep ids in logs and traces where they belong.

What to actually do with this

  • List every alert that fired last month and delete the ones nobody acted on.
  • Pick one service and write down its three signals and one objective.
  • Replace one threshold alert with a burn-rate alert and compare the noise.
  • Find your highest-cardinality metric before your bill does.
This article covers one checkpoint on the roadmap. Open SLOs people actually use on the roadmap → โ€” it lists what this depends on and everything else written about it.

Something wrong or out of date? Open an issue โ€” corrections are welcome and get credited.