Kubernetes bin-packing
NVIDIA MPS/MIG physical GPU virtualization, sub-GPU slicing, and defragmentation packing.
problem it solves
Eliminates idle GPU compute waste from underutilized hardware allocations.
What it does
What it does: Slices physical GPUs into virtual instances and packs workloads to maximize density.
What it watches: GPU VRAM utilization, compute SM occupancy, and pending pod queue depth.
When it triggers: On Kubernetes GPU pod scheduling and workload scaling events.
The Action: Shares physical cards across co-located pods using NVIDIA MPS / MIG partitioning.
How it recovers: Migrates workloads to dedicated cards if resource contention triggers SLA degradation.
What we need from you
- Kubernetes GPU pool with ACE operatorrequired
Requires Kubernetes cluster with NVIDIA GPU Operator and GPUAllocationPolicy CRD.
- MIG-capable hardwarerecommended
Physical GPUs (A100 / H100 / H200) supporting hardware slicing.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Not consulted. Standard K8s scheduler assigns 1 whole GPU per pod. | No k8s_binpacking stage recorded. |
| shadow | Calculates bin-packing efficiency and node consolidation potential. | Stage with action=would_binpack and potential_savings logged. |
| prod | Executes MPS/MIG bin-packing and dynamic node packing on the cluster. | GPU packing ratio and idle node termination logged. |
Current policy
| Virtualization tech | NVIDIA MPS & MIG partitioning | Hardware and software level GPU sharing. |
| Orchestration | K8s GPU Operator (GPUAllocationPolicy) | Automated MPS/MIG geometry application and defragmentation. |
| Target utilization | 85% GPU compute density | Eliminates empty GPU slice waste. |
Worth knowing before you enable it
- ·Software MPS sharing requires application memory isolation validation.
- ·Hardware MIG partitioning requires node reboot or driver re-initialization on geometry change.
- ·Heavy parallel workloads may experience minor CUDA kernel launch contention under high MPS packing.
What it replaces
- ·1-pod-per-GPU waste in standard Kubernetes setups.
- ·Manual static MIG partition scripts.
- ·Over-provisioned idle GPU node cluster costs.
Multi-cluster GPU orchestration and custom scheduling plugins on enterprise.
- ·Cross-cloud Kubernetes cluster GPU federation.
- ·Custom topology-aware GPU scheduling rules.
- ·Automated spot-to-on-demand pod migration.