Renewable CPU-rate budget ('CPU per minute') for long-running guest executions (agents)
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 35/100
- Loại issue
- Tính năng
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Ít trao đổi
- Công nghệ
- javascript, linux, rust, wasm
- Lĩnh vực
- backend, infrastructure, performance
Hướng nghiên cứu
Bắt đầu với V8 CpuBudgetGuard và các giới hạn thực thi được mô tả trong issue, sau đó kiểm tra docs/design/unified-sidecar-runtime.md và crates/execution/src/wasm.rs. Xem xét packages/load-tests/src/compute/server.ts cho kịch bản agent-session. Được xem là hoàn tất khi có các giới hạn CPU có thể gia hạn, xử lý trạng thái nhàn rỗi/không có tiến triển, các giới hạn CPU/WASM đã được di chuyển và các bài kiểm thử cho thấy các session nhàn rỗi hoàn tất trong khi spinner và các lần thực thi bị treo bị dừng.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Problem
Guest execution limits that are totals or wall-clock based are the wrong
shape for long-running, I/O-bound jobs — most importantly agent sessions.
An agent session (e.g. the pi ACP adapter) is a long-lived orchestrator that
runs for minutes across many LLM turns, spending ~95% of that time waiting
on the LLM, the network, and tool I/O — using almost no CPU. A per-execution
wall-clock limit counts that wait time, so a perfectly well-behaved agent is
killed mid-session:
Error: Script execution exceeded the wall-clock limit (limits.jsRuntime.wallClockLimitMs)
ACP adapter process ... exited with code 1 ... live session route evicted
(Observed in the agent load test: with realistic 1.5–6 s LLM latency × N tool
calls, a session runs past the limit and the adapter is terminated.)
Wall-clock conflates three distinct failure modes into one blunt cap:
- using too much CPU (a runaway/spinning execution),
- hung (blocked forever, using ~0 CPU),
- running too long (a legitimately huge session).
For a short untrusted tool call collapsing all three into "kill at 20 s" is
fine. For a long-lived agent runtime it is wrong — you must separate them.
Proposed fix: renewable CPU-rate budget ("CPU per minute")
Bound the long-lived execution by a renewable CPU-rate budget — the Linux
cgroup cpu.max model (quota per period, refilling), i.e. "X CPU-seconds per
1 s window":
- An I/O-bound agent that is mostly idle-on-CPU stays under budget indefinitely
→ runs the whole session, no guillotine. - A runaway/spinning execution pegs a core → burns its per-window allowance →
throttled/killed. This is what wall-clock was trying to catch, done correctly. - Rate-based ⇒ works for arbitrarily long sessions without a magic "max minutes".
Complements:
- Idle / no-progress timeout — a CPU-rate budget can't see a process stuck on
0 CPU (case 2 above), so kill on no LLM/tool activity for N seconds. - Optional session-max ceiling (generous, e.g. 30 min) as a final backstop.
Implementation note: the machinery already exists — the V8 CpuBudgetGuard
(the interrupt that enforces cpuTimeLimitMs today) is a total budget. Making
it renewable (refill over wall-clock time) turns it into exactly this rate
limiter, with no new subsystem and without per-thread cgroups (which
docs/design/unified-sidecar-runtime.md says to avoid for isolates).
Other caps with the same "total, should be renewable" flaw
While here, these per-execution caps are also totals/wall-clock-based and bite
long-running jobs the same way — they should move to the renewable model:
limits.jsRuntime.cpuTimeLimitMs— total accumulated CPU per execution
(defaultDEFAULT_V8_CPU_TIME_LIMIT_MS = 30_000). A long agent doing real work
accumulates >30 s CPU over minutes and is killed. This is the primary target
(wall-clock already defaults to0/disabled; CPU-time does not).limits.resources.maxWasmFuel— WASM fuel, enforced as a wall-clock
timeout for the WASM runtime (crates/execution/src/wasm.rs). Same total/
wall-clock shape; a long-running WASM command should get a renewable fuel rate.
Instantaneous caps (maxProcesses, maxOpenFds, maxSockets, v8HeapLimitMb,
maxWasmMemoryBytes, pendingEvent*, maxFilesystemBytes, maxInodeCount) are
"how many at once / how much stored" — not consumption budgets — and correctly
stay as-is.
Interim state
- Runtime default wall-clock is already disabled (
DEFAULT_V8_WALL_CLOCK_LIMIT_MS = 0). - The load-test actor (
packages/load-tests/src/compute/server.ts) had an explicit
wallClockLimitMs: 20_000that killed agent sessions; it has been set to0
(disabled) with a TODO pointing here. cpuTimeLimitMsis left in place for now; once the renewable budget lands, both
wall-clock and CPU-time should re-enable a sane renewable default.
Acceptance
- Renewable CPU-rate budget (
cpu.max-style) for guest executions, config-exposed. - Idle/no-progress timeout for the hung case.
-
cpuTimeLimitMsandmaxWasmFuelmigrated to (or complemented by) the rate model. - A long (minutes) mostly-idle agent session runs to completion; a CPU spinner
is still throttled/killed; a hung execution is still reaped. - Sane renewable defaults re-enabled (replacing the disabled wall-clock).
- Ngôn ngữ chính
- Rust
- Star
- 4.7k
- Fork
- 263
- Merge trung bình
- 8 giờ 57 phút
- Pull request đã merge (30 ngày)
- 30
Chuẩn bị môi trường
Dự án này không cung cấp dev container, Dockerfile hay hướng dẫn đóng góp, nên bạn cần tự thiết lập môi trường: hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của rivet-dev/agentos
-
Python edits to filesystem.writeFile-created files are reverted by shadow reconciliationCó thể đã có người làm @ankssjain đã nhận 3 ngày trước. Đang mở
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 20/100
rivet-dev/agentos#2022 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 3/5 Nửa ngày Mức phù hợp với người mới 32/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 45/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 66/100
rivet-dev/agentos#1994 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 42/100
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của rivet-dev/agentos
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 75/100
element-hq/lk-jwt-service#248 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
pact-foundation/pact-cli#154 ·
Maintainer thường phản hồi trong vòng 3 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
antithesishq/bombadil#361 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
test(executor_l0): assert execute() TaskOutcome, not only bus events / 断言 execute() 返回的 TaskOutcomeĐang mởtype:debt
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 62/100
skaiy/wild_agentos#425 ·
Maintainer thường phản hồi trong vòng 1 ngày