ob status and execution list do not show an in-flight scheduled job, making an exclusive lock look orphaned
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 48/100
- Tipo de issue
- Error
- Claridad
- Bastante claro
- Estado de actividad
- Activo
- Stack tecnológico
- docker, go
- Área
- cli, devops, infrastructure
Línea de trabajo
Comienza con las implementaciones detrás de ob status, ob execution list y ob schedule list, y luego sigue cómo se registran el estado de los trabajos programados y el bloqueo de rendezvous exclusivo. Reproduce un trabajo exclusivo en curso y compara la salida de cada comando; se considera terminado cuando el trabajo en ejecución, el tiempo que lleva el titular y el estado preciso son visibles sin inspeccionar bloqueos a nivel del host.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
What happens
An operation was refused with:
✗ ob: application scheduling rendezvous remained busy — wait for the current
scheduled job or application operation to finish
while, at the same moment, onebox reported that nothing was running:
$ ob execution list
no durable executions recorded
$ ob status
schedule <weekly-job> last run failed: timeout (exit 143), and nothing has run since ⚠
schedule <5min-job> skipped 20 firings in a row: an application operation is taking its lock ⚠
Host inspection showed the truth: a scheduled job declaring deploy_lock: exclusive had been running for nearly eight hours and legitimately held the rendezvous.
PID 365044 started 10:00:11 elapsed 07:57:46 PPID 1
/bin/sh /etc/systemd/system/ob-<app>-<job>.run
PID 982137 started 15:47:12 docker compose run --rm --no-deps ... <job>
both holding fd 8 on <base>/<app>/schedule.lock. The application lock file itself did not exist, so nothing was stuck — the lock was doing exactly its job.
The bug
An in-flight scheduled job is invisible to ob status and ob execution list.
ob execution listreturnsno durable executions recordedwhile a job has been running for hours.ob statusreports the previous run's outcome and adds "and nothing has run since", which is actively wrong and points away from the truth.
ob schedule list does show the current run's start time as LAST TRIGGER, so the information exists and is simply not joined up.
Why it matters
The refusal message is accurate but gives nothing to act on. Every diagnostic onebox offers then says the system is idle, so the natural conclusion is an orphaned lock from the earlier timeout — that is what I concluded, and it is wrong. The real answer, "a job with an exclusive lock has been running since 10:00 and will hold it until it finishes", is not reachable through any ob command. It took /proc/locks and an fd scan on the host to find it.
The confusion is compounded by skipped being the same word used for a healthy skip, so a schedule that has done nothing for hours looks unremarkable.
Suggested
- Show in-flight scheduled jobs in
ob statusandob execution list. A running exclusive job is the single most useful thing to know when an operation is refused. - Name the holder in the
rendezvous remained busyerror — job name and start time. The data is discoverable from/proc/locksplus an fd scan; surfacing it turns a dead end into a wait. - Do not report "nothing has run since" when something is running.
- Consider whether repeated
skippedoutcomes deserve a distinct signal from ordinary ones.
Environment
- Runner:
ob v2026.9.3 (ffe61080c66b) - Host: Ubuntu, docker 29.1.3, compose 2.40.3
- Jobs: one weekly with
deploy_lock: exclusiveand an 8h timeout, two withdeploy_lock: pinned
- Lenguaje dominante
- Go
- Estrellas
- 3
- Forks
- 0
- Merge medio
- 2 h 46 min
- PR fusionados (30 d)
- 38
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de labstack/onebox
-
bug
Dificultad 1/5 Menos de una hora Aptitud para principiantes 80/100
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 1/5 Menos de una hora Aptitud para principiantes 85/100
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 1/5 Menos de una hora Aptitud para principiantes 85/100
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 3/5 1-2 días Aptitud para principiantes 65/100
Los mantenedores suelen responder en 1 día
-
Dificultad 4/5 3-5 días Aptitud para principiantes 55/100
Los mantenedores suelen responder en 1 día
Todos los issues de labstack/onebox
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
router-for-me/CLIProxyAPI#6399 ·
Los mantenedores suelen responder en 1 día
-
settings.py flaps between reconciles: needsMigrationSetting depends on map iteration orderPosiblemente ocupada @fontaineajulien la tomó hoy. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
pulp/pulp-operator#1691 ·
-
enhancement
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
AOSSIE-Org/DebateAI#611 ·
Los mantenedores suelen responder en 3 días