Workloads and configuration

Deployments, probes and rollouts

Rolling updates that do not drop traffic, and the three probes people keep confusing.

CKADCKA 10 min read

A rolling update is the most routine thing a cluster does and the most common source of self-inflicted downtime. The default settings are reasonable; they are just not the settings most applications need, and the gap shows up only under real traffic.

Three probes, three different jobs

  • startupProbe — "has it finished booting?" While this is failing, the other two are not run at all.
  • readinessProbe — "should it receive traffic?" Failing removes the pod from Service endpoints. It does not restart anything.
  • livenessProbe — "is it wedged?" Failing kills the container.
The classic outage: using a slow livenessProbe as a startup check. An application that takes 90 seconds to warm up, with a livenessProbe that starts checking at 30 seconds, will be killed and restarted forever — and it will never once report an error, because it never got to finish starting. That is what startupProbe exists for.
startupProbe:                 # generous: booting is allowed to be slow
  httpGet: { path: /healthz, port: 8080 }
  failureThreshold: 30
  periodSeconds: 5             # up to 150s to start, checked every 5s

readinessProbe:               # strict: reflects whether dependencies are usable
  httpGet: { path: /ready, port: 8080 }
  periodSeconds: 5

livenessProbe:                # conservative: only true deadlock should trip this
  httpGet: { path: /healthz, port: 8080 }
  periodSeconds: 10
  failureThreshold: 6

A useful rule: /ready may check the database, the cache, the thing you cannot work without. /healthz must check nothing external. If liveness depends on your database, a database blip restarts every pod you own at once, which turns a brief degradation into a cold-start stampede.

maxUnavailable is the dial that matters

spec:
  replicas: 6
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 0        # never go below 6 serving pods
      maxSurge: 2              # briefly run up to 8

The default is 25% of each. With maxUnavailable: 25% and four replicas you are deliberately running on three during every deploy, which is fine until the deploy coincides with a traffic peak. Setting it to 0 with some surge costs you a little headroom and removes a whole category of deploy-time incidents.

Termination is not instant, and the gap bites

When a pod is deleted, two things happen at the same time: the kubelet sends SIGTERM, and the endpoint controller starts removing it from Services. Those are not synchronised. For a second or two, a pod that is already shutting down is still receiving new connections.

lifecycle:
  preStop:
    exec:
      command: ["sleep", "5"]   # keep serving while endpoints propagate
terminationGracePeriodSeconds: 30

A preStop sleep looks crude and is the standard fix. The pod stays up and keeps serving for a few seconds after it is marked for deletion, by which time the proxies have stopped sending it work. Then your application gets its SIGTERM and can drain cleanly.

Watching a rollout properly

kubectl rollout status deploy/my-app --timeout=120s
kubectl rollout history deploy/my-app
kubectl rollout undo deploy/my-app --to-revision=3

# why is it stuck? the Deployment tells you, in conditions
kubectl get deploy my-app -o jsonpath='{.status.conditions[*].message}{"\n"}' 

A stalled rollout usually reports ProgressDeadlineExceeded, which means new pods never became ready within 10 minutes. The Deployment is not the thing to debug at that point — the new pod is.

What to actually do with this

  • Set maxUnavailable: 0 on anything that serves users.
  • Add a startupProbe to anything slow, and make the liveness check depend on nothing external.
  • Add a five-second preStop sleep and watch your deploy-time error rate change.
  • Break a rollout on purpose, then read .status.conditions rather than guessing.
This article covers one checkpoint on the roadmap. Open Deployments, probes and rollouts on the roadmap → — it lists what this depends on and everything else written about it.

Something wrong or out of date? Open an issue — corrections are welcome and get credited.