Troubleshooting under time pressure
A fixed order of operations for a broken cluster, so you stop guessing when it matters most.
Under pressure people debug by hunch, and hunches are biased towards whatever broke last time. A fixed order is slightly slower on the occasions you guess right, and dramatically faster on average.
The order
- Is the pod scheduled?
Pendingis a scheduler problem, never an application problem. - Did the image arrive?
ImagePullBackOffis registry, credentials or a typo. - Did the process start?
CrashLoopBackOffmeans it ran and exited — read the previous logs. - Is it ready? Running but receiving no traffic is a probe or a selector.
- Can it reach what it needs? Now, and only now, is it networking.
Each step rules out a whole class of cause. Skipping to step five is how an afternoon disappears into packet captures for what turns out to be a failing readiness probe.
Events before logs, always
kubectl get events -n prod --sort-by=.lastTimestamp | tail -30
kubectl describe pod my-app-xyz | sed -n '/Events:/,$p'
Events are where the scheduler and kubelet explain themselves: Insufficient cpu, node(s) had untolerated taint, FailedMount, Liveness probe failed. They answer “why is this not running” in a way logs never do, because the application never got far enough to log.
The previous container is the one that failed
kubectl logs my-app-xyz --previous # the instance that crashed
kubectl logs my-app-xyz -c sidecar --previous # and the right container
kubectl get pod my-app-xyz \
-o jsonpath='{.status.containerStatuses[*].lastState.terminated.reason}'
In a crash loop, kubectl logs without --previous shows the fresh container that has not failed yet, which is usually empty. That is the most common reason people believe an application logs nothing on failure.
Exit codes are worth memorising. 137 is SIGKILL — almost always an OOM kill or a failed liveness probe. 143 is SIGTERM, a normal shutdown. 1 or 2 is your application failing on its own terms, so read the logs.
Node NotReady: the first five commands
kubectl get nodes -o wide
kubectl describe node bad-node | sed -n '/Conditions:/,/Addresses:/p'
# then on the node itself
systemctl status kubelet
journalctl -u kubelet -n 100 --no-pager
df -h /var/lib/kubelet && free -m
Node conditions name the problem directly: DiskPressure, MemoryPressure, KubeletNotReady. A full disk on /var/lib/kubelet is the most common cause of a node going NotReady and having its pods evicted, and it is invisible from the cluster side until you look.
Debugging a container with no shell
# attach a debug container to a running pod, sharing its namespaces
kubectl debug -it my-app-xyz --image=nicolaka/netshoot --target=app
# a copy of the pod with the command replaced, when it crashes too fast to exec into
kubectl debug my-app-xyz -it --copy-to=debug-pod --container=app -- sh
Distroless and scratch images have no shell on purpose, which is good for security and awkward at 2am. kubectl debug solves it without rebuilding the image or weakening it.
What to actually do with this
- Write the five-step order somewhere you will see it during an incident.
- Practise
kubectl debugonce, before you need it. - Check disk on your nodes now. It is the cheapest outage to prevent.
Something wrong or out of date? Open an issue โ corrections are welcome and get credited.