[rush] rushd: a client that stops reading its output keeps the batch/lease after its operations finish; every other client blocks and fails with wait-timeout
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 40/100
- Tipo di issue
- Bug
- Chiarezza
- Specificata chiaramente
- Stato di attività
- Attiva
- Stack tecnologico
- nodejs, typescript
Direzione di ricerca
The issue points to the root cause in libraries/rush-daemon/src/PhasedRequestEventSink.ts lines 42-84, specifically the OrderedClientWriter class. Start by examining how flushAsync() is awaited and how batch completion is tied to client output draining. Understand the daemon's request handling and lease management. A fix involves decoupling batch completion from per-client output draining, implementing a bounded buffer, and possibly spilling to a log file. Test with the provided repro steps to verify the fix.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Summary
If one client stops reading its stdout (a pager, Ctrl+S, a suspended or slow agent), the daemon keeps that request's batch, and its execution lease, open until the client has drained all of its output. This continues even after that client's operations have finished. Every other request to the workspace queues behind it and fails with rush-client: daemon admission failed (wait-timeout). after 30 s, even an unrelated no-op build of a different project. The daemon buffers the client-bound frames in memory without bound in the meantime.
Repro steps
Two-project synthetic workspace; p01 prints 100k lines (3.7 MB) with --verbose; RUSH_DAEMON=1:
rush-client rebuild --verbose --only p01 2>&1 | (sleep 45; wc -c) & # client A: its reader stalls for 45 s
sleep 3; rush-client build --only p02 # client B: unrelated no-op, 1.4-1.7 s when run alone
Expected result: A slow consumer only slows itself down. Once its operations have finished, the batch and lease are released and its remaining output drains independently (a bounded buffer, a spill to or replay from the operation log file, or detaching the client after a limit). Unrelated requests never wait on another client's terminal.
Actual result: p01's child process finished by t=16 s, but client B fails with exit 1 at 34.5 s (wait-timeout). With a 30 s stall, B finishes at 32.2 s instead of about 1.5 s. Daemon RSS grows with the buffered output (+10 MB for 15 MB).
Details
Root cause (main @ 60007c9a8c): libraries/rush-daemon/src/PhasedRequestEventSink.ts:42-84 (OrderedClientWriter) chains every event and log chunk onto an unbounded promise tail, and writeLogChunk returns void, so the engine never feels backpressure. flushAsync() is then awaited before the request result and batch completion, so batch completion depends on the slowest client's terminal.
Suggested fix: decouple batch completion from per-client output draining. Release the lease when execution finishes, and drain each client's remaining output on that client's own connection with a bounded buffer. If a limit is exceeded, spill to the log file and tell the client, or disconnect it with an explicit error.
This was found during an automated performance/behavior analysis of rush-client/rushd on Linux (multi-agent scenario) and independently reproduced.
Standard questions
| Question | Answer |
|---|---|
@microsoft/rush globally installed version? |
built from main @ 60007c9a8c (5.179.0) |
rushVersion from rush.json? |
5.179.0 |
pnpmVersion, npmVersion, or yarnVersion from rush.json? |
pnpm@10.27.0 |
(if pnpm) useWorkspaces from pnpm-config.json? |
true |
| Operating system? | Linux (WSL2 Ubuntu 24.04) |
| Would you consider contributing a PR? | Yes |
Node.js version (node -v)? |
22.23.2 |
- Lingua principale
- TypeScript
- Stelle
- 6.5k
- Fork
- 708
- Merge medio
- 4g 13h
- PR unite (30g)
- 62
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di microsoft/rushstack
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 45/100
Tutte le issue di microsoft/rushstack
Issue simili
-
bug(cli): hapi doctor inline-media prints a fabricated B:\ helper-script path in packaged installs Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
-
Crush Aperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 85/100
catppuccin/catppuccin#3125 ·
-
Add a SECURITY.md Aperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 90/100
ElementsProject/cln-application#167 · 1 commento · 1 reazione ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
Quantco/pnpm-licenses#17 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100