[rush] rushd: a client that stops reading its output keeps the batch/lease after its operations finish; every other client blocks and fails with wait-timeout
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 40/100
- Tipo de issue
- Error
- Claridad
- Bien especificado
- Estado de actividad
- Activo
- Stack tecnológico
- nodejs, typescript
Línea de trabajo
The issue points to the root cause in libraries/rush-daemon/src/PhasedRequestEventSink.ts lines 42-84, specifically the OrderedClientWriter class. Start by examining how flushAsync() is awaited and how batch completion is tied to client output draining. Understand the daemon's request handling and lease management. A fix involves decoupling batch completion from per-client output draining, implementing a bounded buffer, and possibly spilling to a log file. Test with the provided repro steps to verify the fix.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Summary
If one client stops reading its stdout (a pager, Ctrl+S, a suspended or slow agent), the daemon keeps that request's batch, and its execution lease, open until the client has drained all of its output. This continues even after that client's operations have finished. Every other request to the workspace queues behind it and fails with rush-client: daemon admission failed (wait-timeout). after 30 s, even an unrelated no-op build of a different project. The daemon buffers the client-bound frames in memory without bound in the meantime.
Repro steps
Two-project synthetic workspace; p01 prints 100k lines (3.7 MB) with --verbose; RUSH_DAEMON=1:
rush-client rebuild --verbose --only p01 2>&1 | (sleep 45; wc -c) & # client A: its reader stalls for 45 s
sleep 3; rush-client build --only p02 # client B: unrelated no-op, 1.4-1.7 s when run alone
Expected result: A slow consumer only slows itself down. Once its operations have finished, the batch and lease are released and its remaining output drains independently (a bounded buffer, a spill to or replay from the operation log file, or detaching the client after a limit). Unrelated requests never wait on another client's terminal.
Actual result: p01's child process finished by t=16 s, but client B fails with exit 1 at 34.5 s (wait-timeout). With a 30 s stall, B finishes at 32.2 s instead of about 1.5 s. Daemon RSS grows with the buffered output (+10 MB for 15 MB).
Details
Root cause (main @ 60007c9a8c): libraries/rush-daemon/src/PhasedRequestEventSink.ts:42-84 (OrderedClientWriter) chains every event and log chunk onto an unbounded promise tail, and writeLogChunk returns void, so the engine never feels backpressure. flushAsync() is then awaited before the request result and batch completion, so batch completion depends on the slowest client's terminal.
Suggested fix: decouple batch completion from per-client output draining. Release the lease when execution finishes, and drain each client's remaining output on that client's own connection with a bounded buffer. If a limit is exceeded, spill to the log file and tell the client, or disconnect it with an explicit error.
This was found during an automated performance/behavior analysis of rush-client/rushd on Linux (multi-agent scenario) and independently reproduced.
Standard questions
| Question | Answer |
|---|---|
@microsoft/rush globally installed version? |
built from main @ 60007c9a8c (5.179.0) |
rushVersion from rush.json? |
5.179.0 |
pnpmVersion, npmVersion, or yarnVersion from rush.json? |
pnpm@10.27.0 |
(if pnpm) useWorkspaces from pnpm-config.json? |
true |
| Operating system? | Linux (WSL2 Ubuntu 24.04) |
| Would you consider contributing a PR? | Yes |
Node.js version (node -v)? |
22.23.2 |
- Lenguaje dominante
- TypeScript
- Estrellas
- 6.5k
- Forks
- 708
- Merge medio
- 4 d 13 h
- PR fusionados (30 d)
- 62
Guía de contribución
No hay ninguna guía de contribución indexada para este repositorio
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de microsoft/rushstack
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
-
Dificultad 4/5 3-5 días Aptitud para principiantes 45/100
Todos los issues de microsoft/rushstack
Issues similares
-
bug(cli): hapi doctor inline-media prints a fabricated B:\ helper-script path in packaged installs Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
-
Crush Abierto
Dificultad 1/5 Menos de una hora Aptitud para principiantes 85/100
catppuccin/catppuccin#3125 ·
-
Add a SECURITY.md Abierto
Dificultad 1/5 Menos de una hora Aptitud para principiantes 90/100
ElementsProject/cln-application#167 · 1 comentario · 1 reacción ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
Quantco/pnpm-licenses#17 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100