Foundations

Containers before Kubernetes

Namespaces, cgroups and image layers by hand, so that pod behaviour stops looking like magic and starts looking like Linux.

KCNA 8 min read

Almost every confusing thing Kubernetes does to a container is something Linux was already doing. If you learn the orchestrator first, you end up memorising behaviour. If you spend an afternoon below it, the same behaviour becomes obvious.

A container is three kernel features and a chroot

There is no container object in the Linux kernel. What you get instead is a process with a restricted view of the system, assembled from namespaces, cgroups and a mounted filesystem. You can build one by hand:

# a new PID, mount, UTS and network namespace, with a shell inside it
sudo unshare --pid --mount --uts --net --fork --mount-proc bash

# inside: you are process 1, and you can see almost nothing
ps aux
hostname container-by-hand
ip addr          # just lo - no route out, because nothing was plugged in yet

That last line is the whole of pod networking in miniature. A fresh network namespace has a loopback interface and nothing else. Something outside has to create a virtual ethernet pair, move one end in, and add routes. On a Kubernetes node, that something is the CNI plugin.

cgroups are why your pod was killed

Namespaces control what a process can see. Control groups control what it can use. On cgroup v2 the interface is a filesystem:

sudo mkdir /sys/fs/cgroup/demo
echo "100M" | sudo tee /sys/fs/cgroup/demo/memory.max
echo $$    | sudo tee /sys/fs/cgroup/demo/cgroup.procs

# now allocate more than that in this shell and watch what happens
cat /sys/fs/cgroup/demo/memory.events     # look at oom_kill
This is the actual mechanism behind OOMKilled. A memory limit in a pod spec becomes memory.max on a cgroup. When the process crosses it, the kernel OOM killer acts — not the kubelet, not the scheduler. That is why the container exits with code 137 and your application logs show nothing: it was not asked to stop.

CPU behaves differently, and the difference matters. A CPU limit becomes a quota per period, so exceeding it does not kill anything — it throttles. A pod that is slow but alive is almost always CPU throttling; a pod that dies abruptly is almost always memory.

Images are layers, and layers are a tax

An image is an ordered stack of tarballs plus a JSON manifest. Each instruction in a Dockerfile that changes the filesystem adds a layer, and layers are immutable — so deleting a file in a later layer hides it without reclaiming the space.

# where the size actually went
docker history --no-trunc --format '{{.Size}}\t{{.CreatedBy}}' your-image:tag | head -20

This stops being an aesthetic concern the moment you work on anything with CUDA or a model in it. A 12 GB image is 12 GB pulled onto every node that has to run it, before your process starts. It is the single most common reason a GPU pod appears to "hang" on first schedule — it is not hanging, it is pulling.

What to actually do with this

  • Run the unshare command above once. Ten minutes, and network policy later will make sense.
  • Put a memory.max on a shell and get OOM-killed deliberately, so you recognise it in production.
  • Run docker history on the largest image you own, and find the one layer that is most of it.
  • Multi-stage builds: compile in one stage, copy only the artefact into a slim final stage. Usually the single biggest win available.

Everything above is a property of Linux, not of Kubernetes. That is the point — it is all still true three abstractions up.

This article covers one checkpoint on the roadmap. Open Containers before Kubernetes on the roadmap → — it lists what this depends on and everything else written about it.

Something wrong or out of date? Open an issue — corrections are welcome and get credited.