[rush] rushd: a client that stops reading its output keeps the batch/lease after its operations finish; every other client blocks and fails with wait-timeout
Personne n'a encore pris cette issue.
Évaluation
- Difficulté
- 4/5
- Temps estimé
- 3-5 jours
- Accessibilité débutants
- 40/100
- Type d'issue
- Bug
- Clarté
- Clairement spécifiée
- Activité
- Active
- Stack technique
- nodejs, typescript
Piste de recherche
The issue points to the root cause in libraries/rush-daemon/src/PhasedRequestEventSink.ts lines 42-84, specifically the OrderedClientWriter class. Start by examining how flushAsync() is awaited and how batch completion is tied to client output draining. Understand the daemon's request handling and lease management. A fix involves decoupling batch completion from per-client output draining, implementing a bounded buffer, and possibly spilling to a log file. Test with the provided repro steps to verify the fix.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Description
Summary
If one client stops reading its stdout (a pager, Ctrl+S, a suspended or slow agent), the daemon keeps that request's batch, and its execution lease, open until the client has drained all of its output. This continues even after that client's operations have finished. Every other request to the workspace queues behind it and fails with rush-client: daemon admission failed (wait-timeout). after 30 s, even an unrelated no-op build of a different project. The daemon buffers the client-bound frames in memory without bound in the meantime.
Repro steps
Two-project synthetic workspace; p01 prints 100k lines (3.7 MB) with --verbose; RUSH_DAEMON=1:
rush-client rebuild --verbose --only p01 2>&1 | (sleep 45; wc -c) & # client A: its reader stalls for 45 s
sleep 3; rush-client build --only p02 # client B: unrelated no-op, 1.4-1.7 s when run alone
Expected result: A slow consumer only slows itself down. Once its operations have finished, the batch and lease are released and its remaining output drains independently (a bounded buffer, a spill to or replay from the operation log file, or detaching the client after a limit). Unrelated requests never wait on another client's terminal.
Actual result: p01's child process finished by t=16 s, but client B fails with exit 1 at 34.5 s (wait-timeout). With a 30 s stall, B finishes at 32.2 s instead of about 1.5 s. Daemon RSS grows with the buffered output (+10 MB for 15 MB).
Details
Root cause (main @ 60007c9a8c): libraries/rush-daemon/src/PhasedRequestEventSink.ts:42-84 (OrderedClientWriter) chains every event and log chunk onto an unbounded promise tail, and writeLogChunk returns void, so the engine never feels backpressure. flushAsync() is then awaited before the request result and batch completion, so batch completion depends on the slowest client's terminal.
Suggested fix: decouple batch completion from per-client output draining. Release the lease when execution finishes, and drain each client's remaining output on that client's own connection with a bounded buffer. If a limit is exceeded, spill to the log file and tell the client, or disconnect it with an explicit error.
This was found during an automated performance/behavior analysis of rush-client/rushd on Linux (multi-agent scenario) and independently reproduced.
Standard questions
| Question | Answer |
|---|---|
@microsoft/rush globally installed version? |
built from main @ 60007c9a8c (5.179.0) |
rushVersion from rush.json? |
5.179.0 |
pnpmVersion, npmVersion, or yarnVersion from rush.json? |
pnpm@10.27.0 |
(if pnpm) useWorkspaces from pnpm-config.json? |
true |
| Operating system? | Linux (WSL2 Ubuntu 24.04) |
| Would you consider contributing a PR? | Yes |
Node.js version (node -v)? |
22.23.2 |
- Langage dominant
- TypeScript
- Étoiles
- 6.5k
- Forks
- 708
- Merge moyen
- 4 j 13 h
- PR mergées (30 j)
- 62
Guide de contribution
Aucun guide de contribution indexé pour ce dépôt
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Autres issues de microsoft/rushstack
-
Difficulté 2/5 1-3 heures Accessibilité débutants 72/100
-
Difficulté 2/5 1-3 heures Accessibilité débutants 65/100
-
Difficulté 2/5 1-3 heures Accessibilité débutants 68/100
-
Difficulté 2/5 1-3 heures Accessibilité débutants 70/100
-
Difficulté 4/5 3-5 jours Accessibilité débutants 45/100
Toutes les issues de microsoft/rushstack
Issues similaires
-
Difficulté 2/5 1-3 heures Accessibilité débutants 84/100
bcgov/bc-wallet-mobile#4761 · 1 commentaire ·
-
external-issue to-triage
Difficulté 2/5 1-3 heures Accessibilité débutants 88/100
-
area-deployment area-integrations triage:bot-seen
Difficulté 2/5 Une demi-journée Accessibilité débutants 86/100
-
Difficulté 2/5 1-3 heures Accessibilité débutants 82/100
-
refactor
Difficulté 2/5 1-3 heures Accessibilité débutants 84/100