[BUG] Stall detection cannot see a session idling inside a tool call: pane-log re-arm keeps the grace alive through `sleep`
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 48/100
- Tipo de issue
- Error
- Claridad
- Bastante claro
- Estado de actividad
- Activo
- Stack tecnológico
- python
- Área
- devtools, observability
Línea de trabajo
Empieza por generic.py, alrededor de _log_activity_key, el bucle de fecha límite de stall cerca de la línea 811 y _sample_weighted_usage cerca de la línea 1154; después inspecciona policy.py para ver la configuración de stall y las superficies de TUI/journal. Compara el crecimiento de la transcripción en vivo con la actividad del pane-log y define cómo deben registrarse y mostrarse los periodos de inactividad sin tratar cada llamada de herramienta larga como un stall; se considera terminado cuando el tramo de inactividad es visible en el run record, el journal y la TUI.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Description
_log_activity_key re-arms dev_stall_grace_s from the tee'd pane log's (mtime_ns, size). A coding CLI repaints its spinner every second while it is blocked inside a tool call, so the pane log grows the whole time. The grace therefore cannot expire while a session sits in sleep, and a session that keeps its turn alive by polling — sleep 590; cat <background-task-output> — is indistinguishable from one streaming a diff.
Nothing in the run record separates the two. nudges stays 0, the heartbeat stays fresh, session-start is the last journal event, and the TUI shows a working session. The only bound is session_timeout_min.
This is the shape #157 named but did not file: "because the session never ends a turn, it never emits a result-less Stop, so dev_stall_grace_s / dev_stall_nudges never engage." #157 asked for prompt teardown once timeout_s fires. This asks for the gap before that: for the hours in between, the orchestrator has no signal at all.
Why sessions do this now. #109 / PR #122 ship CLAUDE_CODE_DISABLE_BACKGROUND_TASKS=1 to stop a session yielding its turn to a background subagent. That removed the sanctioned path but not the incentive: there is still no background-completion re-invocation, so yielding costs a 600 s grace plus a nudge that is itself a submitted turn (dev_stall_nudges_cap, #149). Waiting in-band with sleep is cheaper for the session — and it is the one shape no detector sees. The guard did not remove the wait; it moved the wait somewhere invisible.
Steps to reproduce
- Claude adapter, stock profile (so
CLAUDE_CODE_DISABLE_BACKGROUND_TASKS=1is set),dev_stall_grace_s = 600. - Run a story whose dev session spawns a subagent and then polls for its result rather than doing the work inline — e.g.
sleep 590; cat <task-output-file>, repeated. 590 s is the longest sleep that fits one Bash call, whose timeout caps at 600000 ms. - Watch
bmad-loop tuiand the journal.
Expected behavior
Time a session spends idle inside a tool call is visible. At minimum the run record should let an operator answer "how much of this session's wall clock was spent doing nothing", and the TUI should not present such a session identically to a working one.
Actual behavior
The session reads as healthy for as long as it likes. stall_nudges_sent stays 0 because the grace deadline is pushed forward by every spinner repaint. No journal event marks the idle stretch. session_timeout_min is the only backstop, and it fires on total elapsed time, not on idleness.
Operator-side measurement over two runs on 0.11.0 (Claude adapter, macOS): one review session spent 44.2 of 64.2 minutes in sleep, another 17.0 of 31.8, and one dev session 90 of 107 minutes — 84%. Three single-subagent spawns accounted for 80% of that dev session's sleep; fan-outs amortise one wait across several agents and cost little. Four of seven sessions in the same runs never slept at all, including one that drove 22 subagents through 87 minutes, so this is session improvisation the loop currently permits rather than a deterministic path. Those runs are on a different machine from the checkout these line references were read on, so the numbers are quoted as operator measurements rather than attached artifacts. The run directories still exist and I can pull the journals or a diagnose dump if that would help. The code-level argument above stands on its own either way.
Suggested direction
The adapter already holds what it needs. transcript_path arrives on the first hook event and _sample_weighted_usage already polls that live file mid-turn for token spend. A session inside sleep appends no transcript entries and its weighted usage does not move, while the pane log keeps growing — so the two signals disagree exactly when the session is idle.
Worth separating, because they differ in risk:
- Observability (low risk). Track time since the last transcript append alongside the pane-log key. Surface it in the TUI and journal an event when it crosses a threshold. This alone turns an invisible 90 minutes into a visible one and costs nothing in false stalls.
- Bounding (needs design). Do not simply stall on absent transcript growth — a legitimately long single tool call looks the same (#157's 59-minute
docker runis the counter-example). If a bound is wanted it needs its own knob and a way to tell a blocked tool call from a deliberate poll.
The upstream shape behind all of it — no background-completion re-invocation — is noted for context, not proposed here. As long as it is absent, sessions have a standing incentive to wait in-band, and item 1 at least makes the cost measurable.
Filing this separately from #109 because that issue's resolution is shipped and works for what it targeted; this is the pressure that moved elsewhere afterwards.
Which area is this for?
Orchestrator / control loop
bmad-loop Version
0.11.0
Which coding CLI are you using?
Claude (claude)
Operating System
macOS
Relevant log output
generic.py:971 _log_activity_key -> (st.st_mtime_ns, st.st_size) of logs_dir/<task_id>.log
generic.py:811 key = self._log_activity_key(...); if key != last_activity: stall_deadline = now + grace
generic.py:1154 _sample_weighted_usage(transcript_path, spec) # live transcript already polled mid-turn
policy.py:104 dev_stall_grace_s = 600
policy.py:122 dev_stall_nudges_cap = 6
Confirm
- I've searched for existing issues (#109, #149, #157, #158, #470 are related; none covers in-tool idle detection)
- I'm using the latest version
- Lenguaje dominante
- Python
- Estrellas
- 146
- Forks
- 68
- Merge medio
- 2 d 12 h
- PR fusionados (30 d)
- 41
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de bmad-code-org/bmad-loop
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
bmad-code-org/bmad-loop#835 ·
Los mantenedores suelen responder en 1 día
-
no-artifact with no hint when story specs live in a subfolder (non-recursive artifact glob)Posiblemente ocupada @ahcrm-core la tomó hace 15 días. Abiertoarea:adapters enhancement good first issue P3
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
bmad-code-org/bmad-loop#780 · 2 comentarios ·
Los mantenedores suelen responder en 1 día
-
area:adapters area:psmux bug P4
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
bmad-code-org/bmad-loop#673 · 8 comentarios · 1 reacción ·
Los mantenedores suelen responder en 1 día
-
TUI mirrors a parked permission prompt but offers no way to answer it before Claude Code auto-deniesAbierto
Dificultad 4/5 3-5 días Aptitud para principiantes 55/100
bmad-code-org/bmad-loop#855 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 3/5 1-2 días Aptitud para principiantes 55/100
bmad-code-org/bmad-loop#847 · 2 comentarios · 1 reacción ·
Los mantenedores suelen responder en 1 día
Todos los issues de bmad-code-org/bmad-loop
Issues similares
-
bug llm translation
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Los mantenedores suelen responder en 1 día
-
Arkansas 2025 tax is $1.70 high above $100,000 net taxable income ($3,809 + 3.9% rule)Posiblemente ocupada @PavelMakarchuk la tomó hoy. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 74/100
PolicyEngine/policyengine-us#9828 ·
Los mantenedores suelen responder en 2 días
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
jellyfin/jellyfin-mpv-shim#800 ·
Los mantenedores suelen responder en 1 día
-
skillfs: one malformed chat-log line aborts the entire skill-usage analysis (skill_usage_from_chat_logs.py)Posiblemente ocupada @zjncs la tomó hoy. Abiertocomponent:skillfs
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
agentic-os-org/ANOLISA#6116 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
P4: low query
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
jeffknupp/association#336 ·