Interactive session reclaim
Hibernates idle notebook and remote-IDE kernels to disk and returns their VRAM to the schedulable pool.
problem it solves
Stops forgotten Jupyter and remote-IDE kernels from holding whole GPUs for hours while nobody is running anything on them.
What it does
What it does: Watches interactive kernels for sustained inactivity, warns the owner, then checkpoints the session to disk and releases its VRAM.
What it watches: SM utilization, memory bandwidth, live CUDA contexts and terminal or kernel activity — together, not utilization alone.
When it triggers: After compute stays below the idle floor for longer than the configured window, and only once the warning has gone unanswered.
The Action: Hibernates session state to disk and returns the GPU to the pool. The owner restores in one click with variables intact.
How it recovers: Restore rehydrates the kernel from its checkpoint. Nothing is killed outright, so a wrong call costs a restore, not a day's work.
What we need from you
- Interactive sessions on managed GPU nodesrequired
Jupyter or remote-IDE kernels must run on nodes the control plane schedules; unmanaged sessions are invisible to it.
- Checkpoint storage for session staterequired
Hibernation needs somewhere durable to put the session. Without it the only available action is a kill, which this skill will not do.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Not consulted. Idle interactive kernels hold their GPUs until a human notices. | No interactive_reclaim stage recorded. |
| shadow | Scores sessions and logs which ones it would have hibernated and how much VRAM that would have freed. Nothing is warned, hibernated or released. | Stage with action=would_hibernate_session logged, carrying the idle duration and VRAM held. |
| prod | Warns the owner, then hibernates to disk and releases VRAM once the grace window expires. | Warning, hibernation, VRAM freed and any subsequent restore logged per session. |
Current policy
| Idle floor | Under 1% SM utilization | Matches the IDLE_GPU threshold the utilization classifier already uses, so both loops agree on what idle means. |
| Warning before action | Always, with a configurable grace window | A sweeper that hibernates without warning gets switched off by the first person it surprises. |
| False-positive budget | Under 1 per 1,000 session-hours | Scored against a hand-labelled trace of real sessions, because a warm kernel someone is thinking over is not an idle one. |
Worth knowing before you enable it
- ·Low GPU utilization is not idleness — a researcher reading code with a live kernel looks identical to an abandoned one unless terminal activity is also read.
- ·Hibernation preserves session state, not external side effects: open network connections and child processes do not survive the round trip.
- ·Restore is not instant. A large session pays its checkpoint size back on the way in.
What it replaces
- ·Cron jobs that kill notebooks and lose people's work.
- ·Chasing researchers on chat to free a GPU nobody is using.
- ·Whole 80GB cards held for the lifetime of a prototyping session.
Per-desk idle policy and chargeback-aware grace windows on enterprise.
- ·Different idle windows per desk, so a research pod and a production pod are not held to one threshold.
- ·Grace windows weighted by what the session is costing its owner.
- ·Scheduled reclaim sweeps aligned to a desk's working hours.