Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

ob status and execution list do not show an in-flight scheduled job, making an exclusive lock look orphaned

クローズ
#188 コメント 4 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
48/100
issue の種類
バグ
明瞭さ
おおむね明確
活発さ
活発
技術スタック
docker, go

調査の方向性

ob status、ob execution list、ob schedule list の背後にある実装から始め、スケジュールされたジョブの状態と排他的ランデブーロックがどのように記録されるかを追跡します。実行中の排他的ジョブを再現して各コマンドの出力を比較します。実行中のジョブ、保持時間、正確なステータスがホストレベルのロック検査なしで確認できれば完了です。

索引モデルが issue の本文から書いたものです。

説明

What happens

An operation was refused with:

✗ ob: application scheduling rendezvous remained busy — wait for the current
     scheduled job or application operation to finish

while, at the same moment, onebox reported that nothing was running:

$ ob execution list
no durable executions recorded

$ ob status
schedule <weekly-job>  last run failed: timeout (exit 143), and nothing has run since ⚠
schedule <5min-job>    skipped 20 firings in a row: an application operation is taking its lock ⚠

Host inspection showed the truth: a scheduled job declaring deploy_lock: exclusive had been running for nearly eight hours and legitimately held the rendezvous.

PID 365044  started 10:00:11  elapsed 07:57:46  PPID 1
            /bin/sh /etc/systemd/system/ob-<app>-<job>.run
PID 982137  started 15:47:12  docker compose run --rm --no-deps ... <job>

both holding fd 8 on <base>/<app>/schedule.lock. The application lock file itself did not exist, so nothing was stuck — the lock was doing exactly its job.

The bug

An in-flight scheduled job is invisible to ob status and ob execution list.

  • ob execution list returns no durable executions recorded while a job has been running for hours.
  • ob status reports the previous run's outcome and adds "and nothing has run since", which is actively wrong and points away from the truth.

ob schedule list does show the current run's start time as LAST TRIGGER, so the information exists and is simply not joined up.

Why it matters

The refusal message is accurate but gives nothing to act on. Every diagnostic onebox offers then says the system is idle, so the natural conclusion is an orphaned lock from the earlier timeout — that is what I concluded, and it is wrong. The real answer, "a job with an exclusive lock has been running since 10:00 and will hold it until it finishes", is not reachable through any ob command. It took /proc/locks and an fd scan on the host to find it.

The confusion is compounded by skipped being the same word used for a healthy skip, so a schedule that has done nothing for hours looks unremarkable.

Suggested

  • Show in-flight scheduled jobs in ob status and ob execution list. A running exclusive job is the single most useful thing to know when an operation is refused.
  • Name the holder in the rendezvous remained busy error — job name and start time. The data is discoverable from /proc/locks plus an fd scan; surfacing it turns a dead end into a wait.
  • Do not report "nothing has run since" when something is running.
  • Consider whether repeated skipped outcomes deserve a distinct signal from ordinary ones.

Environment

  • Runner: ob v2026.9.3 (ffe61080c66b)
  • Host: Ubuntu, docker 29.1.3, compose 2.40.3
  • Jobs: one weekly with deploy_lock: exclusive and an 8h timeout, two with deploy_lock: pinned
主要言語
Go
スター
3
フォーク
0
平均マージ
2時間 46分
マージ済み PR(30日)
38

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

labstack/onebox のほかの issue

labstack/onebox の issue をすべて見る

似ている issue

Go の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。