The roadmap says what to learn and in what order. These say how the thing actually behaves, including the parts that only show up once you run it.
Enough of the object model that everything later has somewhere to attach.
Namespaces, cgroups and image layers by hand, so that pod behaviour stops looking like magic and starts looking like Linux.
Why everything is a resource, what the control loop actually does, and how to read any CRD you meet later without waiting for someone to document it.
Getting your own software to run, survive a restart and be configurable without a rebuild.
Rolling updates that do not drop traffic, and the three probes people keep confusing.
ConfigMap reload behaviour, projected volumes, and why your Secret is base64 and not encrypted.
Kustomize overlays and a pull-based reconciler, so environments stop drifting silently and a cluster can be rebuilt from Git alone.
The part where it is your pager. Control plane internals, upgrades, and recovering from your own mistakes.
What each component owns, plus a real backup and restore of etcd on a cluster you can afford to destroy.
Roles that are actually least-privilege, and the three places tenancy leaks anyway.
A fixed order of operations for a broken cluster, so you stop guessing when it matters most.
Where most senior interviews go, and where the CKNE lives. Packets, policy and the mesh question.
Follow a packet from one pod to another across nodes, then do it again with the CNI removed so you can see what it was doing for you.
Default-deny without taking production down, and proving the policy does what you claimed it does.
Ingress to Gateway API route by route, and an honest list of what a service mesh costs you.
Supply chain, admission and runtime. The CKS track, plus the parts no exam covers.
Pod Security Admission first, then a policy engine for the rules PSA cannot express and a webhook only when nothing else will do.
Signing, verifying, and actually failing closed when verification fails — which is the only part that matters.
Seeing a container do something it should not, and having a plan for the next ten minutes rather than inventing one live.
No certification covers this yet, which is exactly why it is worth writing down.
The device plugin model, drivers, and the difference between MIG and time-slicing when cost is the constraint.
vLLM behind a Gateway, with readiness that reflects model load rather than process start — which is where most first attempts go wrong.
Queue-depth autoscaling, cold starts measured in minutes, and the bill as a design constraint rather than an afterthought.
What changes when other teams depend on you and nobody reads your docs.
Requests and limits chosen from data, plus node autoscaling that does not thrash.
Three signals per service, one dashboard, and alerts that correspond to someone being paged.
A CRD and controller with controller-runtime — which is also the shortest honest route to your first upstream contribution.