Home / Software y Cloud / The committee that decides where each container lands

The committee that decides where each container lands

In Kubernetes everything starts with a pod: the smallest unit the system knows how to run, a group of containers that share networking and storage. But a cluster rarely has just one server. It usually has dozens or hundreds of nodes, each with its own CPU, memory and disks. The uncomfortable question is: who decides, and how, on which of those nodes each pod lands?

The answer is a logical component called the kube-scheduler (or simply the “scheduler”). It is not the one that runs your workloads; it decides where they run. Its work starts as soon as you declare you want a new pod: the control server (kube-apiserver) receives the request, stores it in its internal database (etcd) and leaves it “unscheduled”. That is where the scheduler comes in: it watches the queue of orphaned pods and, for each one, runs a two-phase process: filtering and scoring.

The filtering phase: discarding the impossible

The first phase is the most drastic. The scheduler walks through all the nodes of the cluster, applying a series of predicates (hard rules) that remove invalid candidates. These rules rely on the requests and limits you declared when defining the pod: the guaranteed memory and CPU and the maximum it may consume.

If a node has already been assigned too much memory or CPU, it is discarded outright: the scheduler does not over-sell resources. Nodes that fail the placement constraints you requested are also filtered. For example, a nodeSelector forces the node to have a specific label (such as disk=ssd), while affinity and anti-affinity express softer preferences like “place me near frontend pods” or “do not put me next to other pods of this service”. And taints and tolerations act as a repellent: if a node is “tainted” and your pod lacks the matching toleration, that node is vetoed.

The scoring phase: picking the best

Once several valid nodes remain, the scheduler moves to the second phase: scoring. It walks a list of priorities, each returning a score from 0 to 100. A node’s final score is a weighted average of all of them, and the node with the highest score wins.

Priorities follow very different criteria. Resource balancing tends to spread the load so no node becomes saturated. Local density does the opposite: it packs together pods of the same service to exploit cache and avoid network traffic between nodes. Data locality favors the node that already has the image or volume the pod needs on disk, so they do not have to be copied. Deciding which criterion weighs more is a kind of bin-packing with nuances: sometimes tight packing is better, sometimes spreading is, and the outcome depends on the configured weights.

One decision at a time, looking at real state

The scheduler does not decide blindly over a static catalog. It subscribes to the Kubernetes watch API, a mechanism through which it receives continuous change notifications (a node going down, a pod terminating, resources rising or falling). It also queries the resources actually in use on each node, not only what was declared hours ago. That real-time information is what prevents sending two pods together to a node that no longer has room.

Everything ends with a binding call: the scheduler writes to the API that the pod is bound to a specific node, and only then does that node’s kubelet process create and start the real containers.

When things break, the scheduler steps in

The scheduler does not only make decisions when you launch a brand-new pod. If a node breaks or runs out of resources, the pods that ran on it become orphaned, and the controller manager requests fresh copies that the scheduler places again on healthy nodes. That is where its true value shows: it is not a mere initial dispatcher, but the guardian that keeps the cluster running through every failure.

There is even an architecture called multi-scheduler, where you can run several schedulers with different policies and choose, via a field (schedulerName), which one takes charge of each pod. That turns the planner into something like a committee of experts: each with its own rules, but all devoted to a single question — where should each container live.