Renewable CPU-rate budget ('CPU per minute') for long-running guest executions (agents)
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 35/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Tranquilla
- Stack tecnologico
- javascript, linux, rust, wasm
- Ambito
- backend, infrastructure, performance
Direzione di ricerca
Inizia con V8 CpuBudgetGuard e i limiti di esecuzione descritti nell’issue, quindi esamina docs/design/unified-sidecar-runtime.md e crates/execution/src/wasm.rs. Esamina packages/load-tests/src/compute/server.ts per lo scenario agent-session. Il lavoro è completo quando include limiti CPU rinnovabili, gestione dell’inattività/assenza di progressi, limiti CPU/WASM migrati e test che mostrano che le sessioni inattive vengono completate mentre spinner ed esecuzioni bloccate vengono arrestati.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Problem
Guest execution limits that are totals or wall-clock based are the wrong
shape for long-running, I/O-bound jobs — most importantly agent sessions.
An agent session (e.g. the pi ACP adapter) is a long-lived orchestrator that
runs for minutes across many LLM turns, spending ~95% of that time waiting
on the LLM, the network, and tool I/O — using almost no CPU. A per-execution
wall-clock limit counts that wait time, so a perfectly well-behaved agent is
killed mid-session:
Error: Script execution exceeded the wall-clock limit (limits.jsRuntime.wallClockLimitMs)
ACP adapter process ... exited with code 1 ... live session route evicted
(Observed in the agent load test: with realistic 1.5–6 s LLM latency × N tool
calls, a session runs past the limit and the adapter is terminated.)
Wall-clock conflates three distinct failure modes into one blunt cap:
- using too much CPU (a runaway/spinning execution),
- hung (blocked forever, using ~0 CPU),
- running too long (a legitimately huge session).
For a short untrusted tool call collapsing all three into "kill at 20 s" is
fine. For a long-lived agent runtime it is wrong — you must separate them.
Proposed fix: renewable CPU-rate budget ("CPU per minute")
Bound the long-lived execution by a renewable CPU-rate budget — the Linux
cgroup cpu.max model (quota per period, refilling), i.e. "X CPU-seconds per
1 s window":
- An I/O-bound agent that is mostly idle-on-CPU stays under budget indefinitely
→ runs the whole session, no guillotine. - A runaway/spinning execution pegs a core → burns its per-window allowance →
throttled/killed. This is what wall-clock was trying to catch, done correctly. - Rate-based ⇒ works for arbitrarily long sessions without a magic "max minutes".
Complements:
- Idle / no-progress timeout — a CPU-rate budget can't see a process stuck on
0 CPU (case 2 above), so kill on no LLM/tool activity for N seconds. - Optional session-max ceiling (generous, e.g. 30 min) as a final backstop.
Implementation note: the machinery already exists — the V8 CpuBudgetGuard
(the interrupt that enforces cpuTimeLimitMs today) is a total budget. Making
it renewable (refill over wall-clock time) turns it into exactly this rate
limiter, with no new subsystem and without per-thread cgroups (which
docs/design/unified-sidecar-runtime.md says to avoid for isolates).
Other caps with the same "total, should be renewable" flaw
While here, these per-execution caps are also totals/wall-clock-based and bite
long-running jobs the same way — they should move to the renewable model:
limits.jsRuntime.cpuTimeLimitMs— total accumulated CPU per execution
(defaultDEFAULT_V8_CPU_TIME_LIMIT_MS = 30_000). A long agent doing real work
accumulates >30 s CPU over minutes and is killed. This is the primary target
(wall-clock already defaults to0/disabled; CPU-time does not).limits.resources.maxWasmFuel— WASM fuel, enforced as a wall-clock
timeout for the WASM runtime (crates/execution/src/wasm.rs). Same total/
wall-clock shape; a long-running WASM command should get a renewable fuel rate.
Instantaneous caps (maxProcesses, maxOpenFds, maxSockets, v8HeapLimitMb,
maxWasmMemoryBytes, pendingEvent*, maxFilesystemBytes, maxInodeCount) are
"how many at once / how much stored" — not consumption budgets — and correctly
stay as-is.
Interim state
- Runtime default wall-clock is already disabled (
DEFAULT_V8_WALL_CLOCK_LIMIT_MS = 0). - The load-test actor (
packages/load-tests/src/compute/server.ts) had an explicit
wallClockLimitMs: 20_000that killed agent sessions; it has been set to0
(disabled) with a TODO pointing here. cpuTimeLimitMsis left in place for now; once the renewable budget lands, both
wall-clock and CPU-time should re-enable a sane renewable default.
Acceptance
- Renewable CPU-rate budget (
cpu.max-style) for guest executions, config-exposed. - Idle/no-progress timeout for the hung case.
-
cpuTimeLimitMsandmaxWasmFuelmigrated to (or complemented by) the rate model. - A long (minutes) mostly-idle agent session runs to completion; a CPU spinner
is still throttled/killed; a hung execution is still reaped. - Sane renewable defaults re-enabled (replacing the disabled wall-clock).
- Lingua principale
- Rust
- Stelle
- 4.7k
- Fork
- 263
- Merge medio
- 9h 43m
- PR unite (30g)
- 27
Preparare l'ambiente
Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di rivet-dev/agentos
-
Python edits to filesystem.writeFile-created files are reverted by shadow reconciliationForse già presa @ankssjain l’ha presa 1 giorno fa. Aperta
Difficoltà 4/5 3-5 giorni Idoneità per principianti 20/100
rivet-dev/agentos#2022 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 3/5 Mezza giornata Idoneità per principianti 32/100
I maintainer di solito rispondono entro 1 giorno
-
Python launched through guest shell stalls on queued filesystem RPCsForse già presa @ankssjain l’ha presa 7 giorni fa. Aperta
Difficoltà 4/5 3-5 giorni Idoneità per principianti 52/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 45/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 66/100
rivet-dev/agentos#1994 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di rivet-dev/agentos
Issue simili
-
documentation
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 90/100
fastrevmd-lab/rustmistmcp#161 ·
-
arch-audit refactor
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
SocketDev/socket-patch#1011 ·
I maintainer di solito rispondono entro 1 giorno
-
bug user-priority/P2
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 92/100
I maintainer di solito rispondono entro 1 giorno
-
opencode: an unanswered --version probe launches opencode 2 without per-session service isolationAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
I maintainer di solito rispondono entro 1 giorno
-
security-advisory
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
MinBZK/regelrecht#1686 ·
I maintainer di solito rispondono entro 1 giorno