sweep discards a validating result.json unread on any non-completed status; zero-hook-event runs pass run/sweep untouched
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 55/100
Línea de trabajo
Empieza por la gestión de triage y migración de sweep.py; después inspecciona SessionResult.stop_seen, validate_triage, validate_migration y la ruta cmd_validate. Reproduce o prueba una sesión no completada con un result.json validable y una ejecución con cero eventos, y confirma que no se culpe incorrectamente al resultado mientras se haga visible o se pause la falta de conexión del hook.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Observed (0.10.0)
In sweep.py's triage handling:
if result.status != "completed":
plan, errors = None, [session_failure_reason("triage", result)]
else:
plan, errors = validate_triage(result.result_json, open_now)
When the session status is anything but completed, result.result_json is never read. After max_triage_attempts the run escalates with "triage output failed validation: ..." — blaming output that was never validated. The migration path has the same shape (validate_migration runs only on completed). #194's env_fault pause softens the transport-failure case only; a lost event channel is not classified as an env fault.
Incident
Run 20260816-113627-d0b2 (macOS, claude adapter). The hook relay had never been registered in that project, so every session's Stop event was lost and each session read as timeout with session_id: null. Both triage attempts wrote a result.json that passes validate_triage (strict and cache mode alike) — attempt 2 finished its actual work in 59 seconds — and both were discarded on session status alone. The escalation blamed the output; the output was fine. signals.py's own docstring predicts the failure mode: losing Stop events means every session stalls to session_timeout_min — "the loudest possible regression, delivered silently".
Proposals (either or both)
-
Validate the result artifact on non-completed status too. If it validates, either use it — a
timeoutwith a valid, complete artifact is a completed turn whose Stop event was lost — or at minimum attach "result artifact present and passes validation" to the escalation so the operator debugs the event channel rather than the agent.SessionResult.stop_seen(#261) is already the right discriminator:status == "timeout" and not stop_seenwith a validating artifact is the lost-event-channel signature. Routing it like #194's env_fault pause (pause, don't charge attempts) would fit the existing shape. -
Make run/sweep enforce the zero-hook-events condition.
cmd_validatefails on unregistered, missing, unreadable, or stale relays, but nothing stops a run whose sessions produce zero events: it burnssession_timeout_minper session and escalates with the wrong blame. Hard-failing (or pausing) after the first session that ends withstop_seen == Falseand an empty run events dir would surface the miswiring at session one instead of N timeouts later.
Happy to attach the journal/state.json from the incident run.
- Lenguaje dominante
- Python
- Estrellas
- 146
- Forks
- 68
- Merge medio
- 1 d 19 h
- PR fusionados (30 d)
- 44
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de bmad-code-org/bmad-loop
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
bmad-code-org/bmad-loop#835 ·
Los mantenedores suelen responder en 1 día
-
area:adapters area:psmux bug P4
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
bmad-code-org/bmad-loop#673 · 8 comentarios · 1 reacción ·
Los mantenedores suelen responder en 1 día
-
Engine crash on fixable repair after a resolved re-drive that escalated before any spec existedAbierto
Dificultad 4/5 3-5 días Aptitud para principiantes 48/100
bmad-code-org/bmad-loop#860 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 4/5 3-5 días Aptitud para principiantes 55/100
bmad-code-org/bmad-loop#859 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 3/5 1-2 días Aptitud para principiantes 68/100
bmad-code-org/bmad-loop#858 ·
Los mantenedores suelen responder en 1 día
Todos los issues de bmad-code-org/bmad-loop
Issues similares
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 85/100
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 75/100
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 85/100
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 85/100
data-umbrella/du-event-board#225 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100