Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Scheduled-job skip blames an application operation when another scheduled job holds the lock

Cerrado
#183 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
3/5
Tiempo estimado
1-2 días
Aptitud para principiantes
76/100
Tipo de issue
Error
Claridad
Bien especificado
Estado de actividad
Activo
Stack tecnológico
go
Área
devops

Línea de trabajo

Comienza en internal/engine/schedule.go alrededor de skip() en la línea 343 y las comprobaciones de bloqueo en las líneas 356 y 363-364; compara el manejo del bloqueo de aplicación con la ruta de error de schedule.lock y revisa internal/engine/lock.go:96-97 para encontrar el otro poseedor del bloqueo. Traza cómo llega la razón seleccionada al archivo de estado, al notifier y a ob status. La tarea está terminada cuando una colisión entre jobs no se informa como una operación de aplicación y hay cobertura para ambos casos distinguibles.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

bug

What happens

When a scheduled job stands aside because it could not take schedule.lock, it records:

an application operation is taking its lock

That names an application operation — a deploy, a job run, an exec — as the holder. In the common case it is another scheduled job, and no application operation is involved at all.

schedule.lock is taken by both:

  • Application operations, but only around the atomic creation of the application lock (internal/engine/lock.go:96-97), so for milliseconds.
  • Every scheduled job run, for as long as it needs (internal/engine/schedule.go:362) — an exclusive job holds it for its whole run, a pinned one until its release lease exists.

So the message describes the rarer holder and omits the frequent one. An operator reading it is pointed at their own deploys and exec sessions, when nothing they do to those will change the outcome.

The runner already distinguishes the two cases it can name accurately:

  • stand_aside 'another run of this job is still in progress' — the per-job lock (schedule.go:356), accurate.
  • skip 'an application operation holds the deploy lock' — the application lock file, tested directly (schedule.go:364), accurate.

It is only the flock on schedule.lock (schedule.go:363) that guesses, and guesses wrong.

skip() writes the reason it is handed and nothing else (schedule.go:343), so the misleading string is what lands in the state file, the notifier and ob status — there is no better-quality record behind it to fall back on.

Observed

A production application with sixteen scheduled jobs, several sharing cron boundaries (* * * * *, */2 * * * *, two firing together on :00):

schedule sync-sec-submissions skipped 5 firings in a row: an application operation is taking its lock ⚠
schedule sync-bls    skipped 3 firings in a row: an application operation is taking its lock ⚠
schedule sync-cftc-cot skipped 3 firings in a row: an application operation is taking its lock ⚠

Ten of the sixteen showed that reason as their last outcome. This sent me looking at the operator's ob exec sessions, which had run in the same window and were the obvious suspect given the wording. They were not the cause: an exec holds schedule.lock only during application-lock creation, and had it been the application lock the reason would have been the other message. The real contention was job against job.

The diagnosis cost is the point — the message is confident and specific, and it is confidently pointing at the wrong thing.

Expected

The reason names what actually held the lock. The runner cannot see the holder directly, but it does not have to guess either: flock failing here means something holds it, and the two candidates are distinguishable at the point of failure — the application lock file's presence and age is already tested a line later.

Roughly: if the application lock exists and is within its TTL, say so; otherwise say another scheduled job holds it. Failing that, a reason that does not assert a cause it has not established would still be an improvement over one that asserts the wrong one.

Why it matters beyond wording

With catch_up: false a skipped firing is lost rather than deferred, so the count in ob status is missed ingestion, not delayed ingestion. An operator who believes the cause is their own deploy cadence will change something that cannot help, while the real cause — schedule density — goes unexamined.

This is adjacent to #133 but not the same: that is about deploys proceeding during long jobs, this is about a job-versus-job collision being reported as an application operation.

Lenguaje dominante
Go
Estrellas
3
Forks
0
Merge medio
2 h 46 min
PR fusionados (30 d)
38

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de labstack/onebox

Todos los issues de labstack/onebox

Issues similares

Más issues de Go

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.