all docs

Kubernetes bin-packing

NVIDIA MPS/MIG physical GPU virtualization, sub-GPU slicing, and defragmentation packing.

problem it solves
Eliminates idle GPU compute waste from underutilized hardware allocations.

What it does

What it does: Slices physical GPUs into virtual instances and packs workloads to maximize density.

What it watches: GPU VRAM utilization, compute SM occupancy, and pending pod queue depth.

When it triggers: On Kubernetes GPU pod scheduling and workload scaling events.

The Action: Shares physical cards across co-located pods using NVIDIA MPS / MIG partitioning.

How it recovers: Migrates workloads to dedicated cards if resource contention triggers SLA degradation.

What we need from you

  • Kubernetes GPU pool with ACE operatorrequired

    Requires Kubernetes cluster with NVIDIA GPU Operator and GPUAllocationPolicy CRD.

  • MIG-capable hardwarerecommended

    Physical GPUs (A100 / H100 / H200) supporting hardware slicing.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. Standard K8s scheduler assigns 1 whole GPU per pod.No k8s_binpacking stage recorded.
shadowCalculates bin-packing efficiency and node consolidation potential.Stage with action=would_binpack and potential_savings logged.
prodExecutes MPS/MIG bin-packing and dynamic node packing on the cluster.GPU packing ratio and idle node termination logged.

Current policy

Virtualization techNVIDIA MPS & MIG partitioningHardware and software level GPU sharing.
OrchestrationK8s GPU Operator (GPUAllocationPolicy)Automated MPS/MIG geometry application and defragmentation.
Target utilization85% GPU compute densityEliminates empty GPU slice waste.

Worth knowing before you enable it

  • ·Software MPS sharing requires application memory isolation validation.
  • ·Hardware MIG partitioning requires node reboot or driver re-initialization on geometry change.
  • ·Heavy parallel workloads may experience minor CUDA kernel launch contention under high MPS packing.

What it replaces

  • ·1-pod-per-GPU waste in standard Kubernetes setups.
  • ·Manual static MIG partition scripts.
  • ·Over-provisioned idle GPU node cluster costs.

Multi-cluster GPU orchestration and custom scheduling plugins on enterprise.

  • ·Cross-cloud Kubernetes cluster GPU federation.
  • ·Custom topology-aware GPU scheduling rules.
  • ·Automated spot-to-on-demand pod migration.
team@acefleet.dev →