all docs

Interactive session reclaim

Hibernates idle notebook and remote-IDE kernels to disk and returns their VRAM to the schedulable pool.

problem it solves
Stops forgotten Jupyter and remote-IDE kernels from holding whole GPUs for hours while nobody is running anything on them.

What it does

What it does: Watches interactive kernels for sustained inactivity, warns the owner, then checkpoints the session to disk and releases its VRAM.

What it watches: SM utilization, memory bandwidth, live CUDA contexts and terminal or kernel activity — together, not utilization alone.

When it triggers: After compute stays below the idle floor for longer than the configured window, and only once the warning has gone unanswered.

The Action: Hibernates session state to disk and returns the GPU to the pool. The owner restores in one click with variables intact.

How it recovers: Restore rehydrates the kernel from its checkpoint. Nothing is killed outright, so a wrong call costs a restore, not a day's work.

What we need from you

  • Interactive sessions on managed GPU nodesrequired

    Jupyter or remote-IDE kernels must run on nodes the control plane schedules; unmanaged sessions are invisible to it.

  • Checkpoint storage for session staterequired

    Hibernation needs somewhere durable to put the session. Without it the only available action is a kill, which this skill will not do.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. Idle interactive kernels hold their GPUs until a human notices.No interactive_reclaim stage recorded.
shadowScores sessions and logs which ones it would have hibernated and how much VRAM that would have freed. Nothing is warned, hibernated or released.Stage with action=would_hibernate_session logged, carrying the idle duration and VRAM held.
prodWarns the owner, then hibernates to disk and releases VRAM once the grace window expires.Warning, hibernation, VRAM freed and any subsequent restore logged per session.

Current policy

Idle floorUnder 1% SM utilizationMatches the IDLE_GPU threshold the utilization classifier already uses, so both loops agree on what idle means.
Warning before actionAlways, with a configurable grace windowA sweeper that hibernates without warning gets switched off by the first person it surprises.
False-positive budgetUnder 1 per 1,000 session-hoursScored against a hand-labelled trace of real sessions, because a warm kernel someone is thinking over is not an idle one.

Worth knowing before you enable it

  • ·Low GPU utilization is not idleness — a researcher reading code with a live kernel looks identical to an abandoned one unless terminal activity is also read.
  • ·Hibernation preserves session state, not external side effects: open network connections and child processes do not survive the round trip.
  • ·Restore is not instant. A large session pays its checkpoint size back on the way in.

What it replaces

  • ·Cron jobs that kill notebooks and lose people's work.
  • ·Chasing researchers on chat to free a GPU nobody is using.
  • ·Whole 80GB cards held for the lifetime of a prototyping session.

Per-desk idle policy and chargeback-aware grace windows on enterprise.

  • ·Different idle windows per desk, so a research pod and a production pod are not held to one threshold.
  • ·Grace windows weighted by what the session is costing its owner.
  • ·Scheduled reclaim sweeps aligned to a desk's working hours.
team@acefleet.dev →