[rush] rushd: a client that stops reading its output keeps the batch/lease after its operations finish; every other client blocks and fails with wait-timeout
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 40/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Đặc tả rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- nodejs, typescript
Hướng nghiên cứu
The issue points to the root cause in libraries/rush-daemon/src/PhasedRequestEventSink.ts lines 42-84, specifically the OrderedClientWriter class. Start by examining how flushAsync() is awaited and how batch completion is tied to client output draining. Understand the daemon's request handling and lease management. A fix involves decoupling batch completion from per-client output draining, implementing a bounded buffer, and possibly spilling to a log file. Test with the provided repro steps to verify the fix.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Summary
If one client stops reading its stdout (a pager, Ctrl+S, a suspended or slow agent), the daemon keeps that request's batch, and its execution lease, open until the client has drained all of its output. This continues even after that client's operations have finished. Every other request to the workspace queues behind it and fails with rush-client: daemon admission failed (wait-timeout). after 30 s, even an unrelated no-op build of a different project. The daemon buffers the client-bound frames in memory without bound in the meantime.
Repro steps
Two-project synthetic workspace; p01 prints 100k lines (3.7 MB) with --verbose; RUSH_DAEMON=1:
rush-client rebuild --verbose --only p01 2>&1 | (sleep 45; wc -c) & # client A: its reader stalls for 45 s
sleep 3; rush-client build --only p02 # client B: unrelated no-op, 1.4-1.7 s when run alone
Expected result: A slow consumer only slows itself down. Once its operations have finished, the batch and lease are released and its remaining output drains independently (a bounded buffer, a spill to or replay from the operation log file, or detaching the client after a limit). Unrelated requests never wait on another client's terminal.
Actual result: p01's child process finished by t=16 s, but client B fails with exit 1 at 34.5 s (wait-timeout). With a 30 s stall, B finishes at 32.2 s instead of about 1.5 s. Daemon RSS grows with the buffered output (+10 MB for 15 MB).
Details
Root cause (main @ 60007c9a8c): libraries/rush-daemon/src/PhasedRequestEventSink.ts:42-84 (OrderedClientWriter) chains every event and log chunk onto an unbounded promise tail, and writeLogChunk returns void, so the engine never feels backpressure. flushAsync() is then awaited before the request result and batch completion, so batch completion depends on the slowest client's terminal.
Suggested fix: decouple batch completion from per-client output draining. Release the lease when execution finishes, and drain each client's remaining output on that client's own connection with a bounded buffer. If a limit is exceeded, spill to the log file and tell the client, or disconnect it with an explicit error.
This was found during an automated performance/behavior analysis of rush-client/rushd on Linux (multi-agent scenario) and independently reproduced.
Standard questions
| Question | Answer |
|---|---|
@microsoft/rush globally installed version? |
built from main @ 60007c9a8c (5.179.0) |
rushVersion from rush.json? |
5.179.0 |
pnpmVersion, npmVersion, or yarnVersion from rush.json? |
pnpm@10.27.0 |
(if pnpm) useWorkspaces from pnpm-config.json? |
true |
| Operating system? | Linux (WSL2 Ubuntu 24.04) |
| Would you consider contributing a PR? | Yes |
Node.js version (node -v)? |
22.23.2 |
- Ngôn ngữ chính
- TypeScript
- Star
- 6.5k
- Fork
- 708
- Merge trung bình
- 4 ngày 13 giờ
- Pull request đã merge (30 ngày)
- 62
Hướng dẫn đóng góp
Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của microsoft/rushstack
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 45/100
Tất cả issue của microsoft/rushstack
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
bcgov/bc-wallet-mobile#4761 · 1 bình luận ·
-
external-issue to-triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
-
area-deployment area-integrations triage:bot-seen
Độ khó 2/5 Nửa ngày Mức phù hợp với người mới 86/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
-
refactor
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100