Scheduled-job skip blames an application operation when another scheduled job holds the lock
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 3/5
- Tiempo estimado
- 1-2 días
- Aptitud para principiantes
- 76/100
Línea de trabajo
Comienza en internal/engine/schedule.go alrededor de skip() en la línea 343 y las comprobaciones de bloqueo en las líneas 356 y 363-364; compara el manejo del bloqueo de aplicación con la ruta de error de schedule.lock y revisa internal/engine/lock.go:96-97 para encontrar el otro poseedor del bloqueo. Traza cómo llega la razón seleccionada al archivo de estado, al notifier y a ob status. La tarea está terminada cuando una colisión entre jobs no se informa como una operación de aplicación y hay cobertura para ambos casos distinguibles.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
What happens
When a scheduled job stands aside because it could not take schedule.lock, it records:
an application operation is taking its lock
That names an application operation — a deploy, a job run, an exec — as the holder. In the common case it is another scheduled job, and no application operation is involved at all.
schedule.lock is taken by both:
- Application operations, but only around the atomic creation of the application lock (
internal/engine/lock.go:96-97), so for milliseconds. - Every scheduled job run, for as long as it needs (
internal/engine/schedule.go:362) — an exclusive job holds it for its whole run, a pinned one until its release lease exists.
So the message describes the rarer holder and omits the frequent one. An operator reading it is pointed at their own deploys and exec sessions, when nothing they do to those will change the outcome.
The runner already distinguishes the two cases it can name accurately:
stand_aside 'another run of this job is still in progress'— the per-job lock (schedule.go:356), accurate.skip 'an application operation holds the deploy lock'— the application lock file, tested directly (schedule.go:364), accurate.
It is only the flock on schedule.lock (schedule.go:363) that guesses, and guesses wrong.
skip() writes the reason it is handed and nothing else (schedule.go:343), so the misleading string is what lands in the state file, the notifier and ob status — there is no better-quality record behind it to fall back on.
Observed
A production application with sixteen scheduled jobs, several sharing cron boundaries (* * * * *, */2 * * * *, two firing together on :00):
schedule sync-sec-submissions skipped 5 firings in a row: an application operation is taking its lock ⚠
schedule sync-bls skipped 3 firings in a row: an application operation is taking its lock ⚠
schedule sync-cftc-cot skipped 3 firings in a row: an application operation is taking its lock ⚠
Ten of the sixteen showed that reason as their last outcome. This sent me looking at the operator's ob exec sessions, which had run in the same window and were the obvious suspect given the wording. They were not the cause: an exec holds schedule.lock only during application-lock creation, and had it been the application lock the reason would have been the other message. The real contention was job against job.
The diagnosis cost is the point — the message is confident and specific, and it is confidently pointing at the wrong thing.
Expected
The reason names what actually held the lock. The runner cannot see the holder directly, but it does not have to guess either: flock failing here means something holds it, and the two candidates are distinguishable at the point of failure — the application lock file's presence and age is already tested a line later.
Roughly: if the application lock exists and is within its TTL, say so; otherwise say another scheduled job holds it. Failing that, a reason that does not assert a cause it has not established would still be an improvement over one that asserts the wrong one.
Why it matters beyond wording
With catch_up: false a skipped firing is lost rather than deferred, so the count in ob status is missed ingestion, not delayed ingestion. An operator who believes the cause is their own deploy cadence will change something that cannot help, while the real cause — schedule density — goes unexamined.
This is adjacent to #133 but not the same: that is about deploys proceeding during long jobs, this is about a job-versus-job collision being reported as an application operation.
- Lenguaje dominante
- Go
- Estrellas
- 3
- Forks
- 0
- Merge medio
- 2 h 46 min
- PR fusionados (30 d)
- 38
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de labstack/onebox
-
bug
Dificultad 1/5 Menos de una hora Aptitud para principiantes 80/100
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 1/5 Menos de una hora Aptitud para principiantes 85/100
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 1/5 Menos de una hora Aptitud para principiantes 85/100
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 3/5 1-2 días Aptitud para principiantes 65/100
Los mantenedores suelen responder en 1 día
-
Dificultad 4/5 3-5 días Aptitud para principiantes 55/100
Los mantenedores suelen responder en 1 día
Todos los issues de labstack/onebox
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
router-for-me/CLIProxyAPI#6399 ·
Los mantenedores suelen responder en 1 día
-
settings.py flaps between reconciles: needsMigrationSetting depends on map iteration orderPosiblemente ocupada @fontaineajulien la tomó hoy. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
pulp/pulp-operator#1691 ·
-
enhancement
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
AOSSIE-Org/DebateAI#611 ·
Los mantenedores suelen responder en 3 días