Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Renewable CPU-rate budget ('CPU per minute') for long-running guest executions (agents)

オープン
#1,817 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
35/100
issue の種類
機能追加
明瞭さ
おおむね明確
活発さ
静か
技術スタック
javascript, linux, rust, wasm

調査の方向性

V8 CpuBudgetGuard と issue に記載されている実行制限から始め、次に docs/design/unified-sidecar-runtime.md と crates/execution/src/wasm.rs を調べてください。agent-session シナリオについて packages/load-tests/src/compute/server.ts を確認してください。完了条件は、更新可能な CPU 制限、アイドル状態/進捗なしの処理、移行された CPU/WASM 上限、およびアイドル状態のセッションは完了する一方で spinner とハングした実行は停止されることを示すテストです。

索引モデルが issue の本文から書いたものです。

説明

Problem

Guest execution limits that are totals or wall-clock based are the wrong
shape for long-running, I/O-bound jobs — most importantly agent sessions.

An agent session (e.g. the pi ACP adapter) is a long-lived orchestrator that
runs for minutes across many LLM turns, spending ~95% of that time waiting
on the LLM, the network, and tool I/O — using almost no CPU. A per-execution
wall-clock limit counts that wait time, so a perfectly well-behaved agent is
killed mid-session:

Error: Script execution exceeded the wall-clock limit (limits.jsRuntime.wallClockLimitMs)
ACP adapter process ... exited with code 1 ... live session route evicted

(Observed in the agent load test: with realistic 1.5–6 s LLM latency × N tool
calls, a session runs past the limit and the adapter is terminated.)

Wall-clock conflates three distinct failure modes into one blunt cap:

  1. using too much CPU (a runaway/spinning execution),
  2. hung (blocked forever, using ~0 CPU),
  3. running too long (a legitimately huge session).

For a short untrusted tool call collapsing all three into "kill at 20 s" is
fine. For a long-lived agent runtime it is wrong — you must separate them.

Proposed fix: renewable CPU-rate budget ("CPU per minute")

Bound the long-lived execution by a renewable CPU-rate budget — the Linux
cgroup cpu.max model (quota per period, refilling), i.e. "X CPU-seconds per
1 s window":

  • An I/O-bound agent that is mostly idle-on-CPU stays under budget indefinitely
    → runs the whole session, no guillotine.
  • A runaway/spinning execution pegs a core → burns its per-window allowance →
    throttled/killed. This is what wall-clock was trying to catch, done correctly.
  • Rate-based ⇒ works for arbitrarily long sessions without a magic "max minutes".

Complements:

  • Idle / no-progress timeout — a CPU-rate budget can't see a process stuck on
    0 CPU (case 2 above), so kill on no LLM/tool activity for N seconds.
  • Optional session-max ceiling (generous, e.g. 30 min) as a final backstop.

Implementation note: the machinery already exists — the V8 CpuBudgetGuard
(the interrupt that enforces cpuTimeLimitMs today) is a total budget. Making
it renewable (refill over wall-clock time) turns it into exactly this rate
limiter, with no new subsystem and without per-thread cgroups (which
docs/design/unified-sidecar-runtime.md says to avoid for isolates).

Other caps with the same "total, should be renewable" flaw

While here, these per-execution caps are also totals/wall-clock-based and bite
long-running jobs the same way — they should move to the renewable model:

  • limits.jsRuntime.cpuTimeLimitMs — total accumulated CPU per execution
    (default DEFAULT_V8_CPU_TIME_LIMIT_MS = 30_000). A long agent doing real work
    accumulates >30 s CPU over minutes and is killed. This is the primary target
    (wall-clock already defaults to 0/disabled; CPU-time does not).
  • limits.resources.maxWasmFuel — WASM fuel, enforced as a wall-clock
    timeout for the WASM runtime (crates/execution/src/wasm.rs). Same total/
    wall-clock shape; a long-running WASM command should get a renewable fuel rate.

Instantaneous caps (maxProcesses, maxOpenFds, maxSockets, v8HeapLimitMb,
maxWasmMemoryBytes, pendingEvent*, maxFilesystemBytes, maxInodeCount) are
"how many at once / how much stored" — not consumption budgets — and correctly
stay as-is.

Interim state

  • Runtime default wall-clock is already disabled (DEFAULT_V8_WALL_CLOCK_LIMIT_MS = 0).
  • The load-test actor (packages/load-tests/src/compute/server.ts) had an explicit
    wallClockLimitMs: 20_000 that killed agent sessions; it has been set to 0
    (disabled) with a TODO pointing here.
  • cpuTimeLimitMs is left in place for now; once the renewable budget lands, both
    wall-clock and CPU-time should re-enable a sane renewable default.

Acceptance

  • Renewable CPU-rate budget (cpu.max-style) for guest executions, config-exposed.
  • Idle/no-progress timeout for the hung case.
  • cpuTimeLimitMs and maxWasmFuel migrated to (or complemented by) the rate model.
  • A long (minutes) mostly-idle agent session runs to completion; a CPU spinner
    is still throttled/killed; a hung execution is still reaped.
  • Sane renewable defaults re-enabled (replacing the disabled wall-clock).
主要言語
Rust
スター
4.7k
フォーク
263
平均マージ
6時間 32分
マージ済み PR(30日)
26

環境構築

このプロジェクトには開発コンテナ、Dockerfile、コントリビューションガイドがありません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

rivet-dev/agentos のほかの issue

rivet-dev/agentos の issue をすべて見る

似ている issue

Rust の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。