all docs

NVMe weight cache

Keeps weights, dataset partitions and container layers on node-local NVMe so scale-up stops re-pulling them.

problem it solves
Stops every new replica from dragging the same multi-gigabyte weights across the storage fabric before it can serve a token.

What it does

What it does: Maintains a content-addressed cache on each node's local NVMe of model weights, frequently-read dataset partitions and base container layers.

What it watches: Cache hit rate per node, bytes pulled from shared storage, and time-to-first-token on a cold replica.

When it triggers: On every pull. A hit is served from local NVMe; a miss falls through to the shared fabric and is then admitted to the cache.

The Action: Uses scheduler lookahead to pre-warm the node a job is about to land on, rather than caching reactively after the miss has already been paid.

How it recovers: A corrupt or evicted entry falls through to the shared fabric transparently — the cache is never the only copy.

What we need from you

  • Local NVMe on GPU nodesrequired

    Nodes must expose a writable local NVMe mount sized for the working set. Network-attached storage defeats the purpose.

  • Fleet nodes reporting telemetryrequired

    The node agent supplies the cache hit rate and pull volume this skill is scored on.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. Every cold start pulls the full artifact set across the shared storage fabric.No nvme_weight_cache stage recorded.
shadowRecords what a cache would have hit and the bytes it would have saved, without writing to local disk.Stage with action=would_cache_artifact logged, carrying the counterfactual hit rate.
prodServes artifacts from local NVMe and pre-warms nodes ahead of placement.Hit rate, bytes served locally and cold-start delta logged per node.

Current policy

Admission policyContent-addressed, admit on second referenceA one-off pull is not cached; the cache is for artifacts a fleet reads repeatedly.
EvictionLeast-recently-used, scheduler-awareAn artifact a pending job is about to need is not evicted to make room for one nothing has asked for.
Target hit rate90% on a repeated jobMeasured on replayed real traffic, not a synthetic benchmark — a synthetic loop trivially hits 100%.

Worth knowing before you enable it

  • ·The first pull of any artifact is always a miss; the win shows up on the second and later replicas, not the first.
  • ·Cache sizing is a working-set question, not a total-corpus one — a node does not need every model your fleet can serve.
  • ·Pre-warming depends on scheduler lookahead, so it degrades to reactive caching wherever placement is decided outside the control plane.

What it replaces

  • ·Hand-rolled rsync pre-warm scripts in node bootstrap.
  • ·Storage fabric saturation alarms every time the fleet scales up.
  • ·Baking model weights into container images to dodge the pull.

Cross-rack peer fetch and dataset-aware pre-warming on enterprise.

  • ·Peer-to-peer cache fill from a neighbouring node instead of the shared fabric.
  • ·Market-data partition pre-warming driven by the job's declared date range.
  • ·Per-desk cache residency guarantees for latency-sensitive artifacts.
team@acefleet.dev →