Kubernetes GPU scheduling with node pools comes down to one contract: a vendor device plugin registers the accelerators with the kubelet, the node advertises a schedulable extended resource such as nvidia.com/gpu, and your containers consume it through resource limits. The scheduler then treats a GPU like any other allocatable resource. What turns that mechanism into a production platform is the node pool: a group of identically configured GPU nodes that you can label, taint, scale and upgrade as a unit. Pool design decides whether expensive accelerators sit idle next to CPU-only pods, which is the largest avoidable cost in most GPU estates — a problem we examined from the serving side in our guide to cutting inference spend with model routing.
How GPU scheduling works
Under the device plugin model, described in the official Schedule GPUs documentation, an administrator installs vendor GPU drivers on the nodes and runs the vendor device plugin as a DaemonSet. After the plugin registers with the kubelet, the node status advertises the new extended resource, and pods consume GPUs by requesting that resource like cpu or memory. The rules are stricter than for ordinary resources: GPUs may only be specified in the limits section, requests and limits must be equal when both are present, and a GPU request without a limit is rejected. Fractional GPUs are not addressable at all — quantities are whole numbers — so time-slicing and MIG partitioning happen inside the plugin, not in the scheduler.
When different nodes carry different accelerator types, placement becomes a labelling problem. Label each node with its accelerator type and memory class, then target pods with nodeSelector or node affinity. Because the extended resource only exists on nodes where the plugin ran, an unsatisfiable combination fails visibly as a Pending pod rather than a silent performance cliff.
Designing GPU node pools
Keep each pool homogeneous: one machine family, one accelerator type, one driver version. Heterogeneous pools force every workload to reason about per-node differences and break the assumption that any node in the pool is interchangeable. Homogeneity also keeps the driver matrix testable — a plugin that works against driver branch A on H100 nodes says nothing about L40S nodes in the same pool.
Node pool boundaries are also drain and upgrade boundaries. When a Kubernetes upgrade drains a pool of eight-GPU nodes, that is real serving capacity leaving the rotation for minutes at a time, so size pools such that one can drain while your throughput SLO holds. Use a stable label scheme — accelerator type, GPU memory, driver branch — and avoid labels that encode volatile facts such as free memory, which belong in monitoring, not scheduling metadata.
Taints, labels, and scheduling
Taints keep CPU workloads off accelerators and keep GPU workloads from landing on CPU-only nodes. On GKE, part of this is automatic: GKE automatically taints GPU nodes with nvidia.com/gpu=present:NoSchedule when you add a GPU node pool to a cluster that already runs a non-GPU pool, and the ExtendedResourceToleration admission controller injects tolerations into pods that actually request GPUs. The practical effect is that GPU nodes scale down quickly when demand disappears, because nothing else can park on them. On self-managed clusters, apply the same pattern by hand: taint the pool at creation with a NoSchedule effect, tolerate the taint in the workload, and pin placement with node affinity rather than nodeName, which bypasses the scheduler entirely. The vendor-side mechanics are documented in the GKE Standard GPU node pool guide.
Manual labelling rots as hardware churns. Node Feature Discovery (NFD) detects hardware features available on each node and advertises them as node labels, extended resources, annotations and taints, so a rebuilt or replaced GPU node inherits the correct scheduling metadata without a human relabelling step. Deploy NFD before you accumulate your second accelerator generation.
Autoscaling GPU capacity
The cluster autoscaler scales a node pool when pending pods fit nowhere that already exists. Combined with taints and tolerations, this lets GPU pools sit at zero nodes during quiet periods and expand on demand. Budget for reality on the way up: GPU node provisioning includes driver initialization and multi-gigabyte image pulls, so scale-up latency is measured in minutes, not seconds, and headroom must be planned rather than discovered. On GKE, node auto-provisioning can create GPU pools from pod requirements directly, with cluster-wide resource limits acting as the ceiling. Spot VMs cut accelerator prices substantially at the cost of frequent disruptions, which pairs well with checkpointed training jobs and poorly with latency-bound inference.
DRA versus device plugins
Dynamic Resource Allocation (DRA) replaces the integer-in-limits contract with a claim-based flow: platform teams define device classes, workloads reference ResourceClaimTemplates, and drivers publish per-node inventories. In this flow, the Kubernetes scheduler reads ResourceSlices published by DRA drivers to decide which devices to allocate, and the kubelet then attaches the allocated hardware to the pod through the driver. Operators filter devices by attributes — GPU memory, vendor, model — instead of hand-selecting nodes, and, as the GKE DRA documentation notes, the platform can place accelerated pods based on device availability without the app operator knowing node configurations.
| Dimension | Device plugins | DRA |
|---|---|---|
| Request style | Integer extended resource in limits | ResourceClaim with attribute filters |
| Node selection | Manual labels, selectors, taints | Scheduler-driven from ResourceSlices |
| Sharing and partitioning | Vendor time-slicing or MIG outside the scheduler | Structured allocations declared by drivers |
| Maturity | Stable since Kubernetes v1.26 | Newer; verify per-platform driver support |
Migrating a fleet is not urgent: device plugins remain the stable default, and DRA adoption should be gated the same way you gate any infrastructure change — with evidence, not enthusiasm. If you already run model-quality gates in your delivery pipeline, as in our walkthrough of LLM evaluation pipelines in CI, extend the same discipline to scheduler migrations.
A practical rollout checklist
- Create one homogeneous GPU node pool per accelerator type and taint it NoSchedule at creation.
- Install vendor drivers through a DaemonSet or platform management, and pin the driver version per pool.
- Deploy Node Feature Discovery so accelerator labels survive node replacement.
- Express GPU demand as whole-number limits only, with tolerations and node affinity for the target pool.
- Enable autoscaling on GPU pools with a minimum of zero and confirm scale-down works with toleration semantics.
- Alert on pending GPU pods and per-pool utilization, and drain one pool at a time during upgrades.
Done in this order, each step is independently reversible, and the scheduler’s behaviour stays explainable when a GPU pod is stuck Pending at three in the morning.