NVMe weight cache
Keeps weights, dataset partitions and container layers on node-local NVMe so scale-up stops re-pulling them.
problem it solves
Stops every new replica from dragging the same multi-gigabyte weights across the storage fabric before it can serve a token.
What it does
What it does: Maintains a content-addressed cache on each node's local NVMe of model weights, frequently-read dataset partitions and base container layers.
What it watches: Cache hit rate per node, bytes pulled from shared storage, and time-to-first-token on a cold replica.
When it triggers: On every pull. A hit is served from local NVMe; a miss falls through to the shared fabric and is then admitted to the cache.
The Action: Uses scheduler lookahead to pre-warm the node a job is about to land on, rather than caching reactively after the miss has already been paid.
How it recovers: A corrupt or evicted entry falls through to the shared fabric transparently — the cache is never the only copy.
What we need from you
- Local NVMe on GPU nodesrequired
Nodes must expose a writable local NVMe mount sized for the working set. Network-attached storage defeats the purpose.
- Fleet nodes reporting telemetryrequired
The node agent supplies the cache hit rate and pull volume this skill is scored on.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Not consulted. Every cold start pulls the full artifact set across the shared storage fabric. | No nvme_weight_cache stage recorded. |
| shadow | Records what a cache would have hit and the bytes it would have saved, without writing to local disk. | Stage with action=would_cache_artifact logged, carrying the counterfactual hit rate. |
| prod | Serves artifacts from local NVMe and pre-warms nodes ahead of placement. | Hit rate, bytes served locally and cold-start delta logged per node. |
Current policy
| Admission policy | Content-addressed, admit on second reference | A one-off pull is not cached; the cache is for artifacts a fleet reads repeatedly. |
| Eviction | Least-recently-used, scheduler-aware | An artifact a pending job is about to need is not evicted to make room for one nothing has asked for. |
| Target hit rate | 90% on a repeated job | Measured on replayed real traffic, not a synthetic benchmark — a synthetic loop trivially hits 100%. |
Worth knowing before you enable it
- ·The first pull of any artifact is always a miss; the win shows up on the second and later replicas, not the first.
- ·Cache sizing is a working-set question, not a total-corpus one — a node does not need every model your fleet can serve.
- ·Pre-warming depends on scheduler lookahead, so it degrades to reactive caching wherever placement is decided outside the control plane.
What it replaces
- ·Hand-rolled rsync pre-warm scripts in node bootstrap.
- ·Storage fabric saturation alarms every time the fleet scales up.
- ·Baking model weights into container images to dodge the pull.
Cross-rack peer fetch and dataset-aware pre-warming on enterprise.
- ·Peer-to-peer cache fill from a neighbouring node instead of the shared fabric.
- ·Market-data partition pre-warming driven by the job's declared date range.
- ·Per-desk cache residency guarantees for latency-sensitive artifacts.