worker: stop awaiting each claim batch, so concurrency limits and /health mean something
メンテナーはふだん 7 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 35/100
- issue の種類
- バグ
- 明瞭さ
- 説明が足りない
- 活発さ
- 活発
- 技術スタック
- typescript
調査の方向性
まず poll()、schedulePoll()、buildHealthResponse、claimAndProcessJobs、job-cleanup.ts を追跡し、次に runWithTimeout、JobHandlerContext、stop() を調べます。変更は、ポーリング、ヘルス、失敗したジョブの処理、シャットダウン、外部副作用について、それぞれ分離した関心事として定義します。完了の条件は、並行実行数の計上とヘルスが正確なまま維持され、失敗、タイムアウト、シャットダウン時の動作が安全かつ有界であることです。
索引モデルが issue の本文から書いたものです。
説明
From the review of #224. These are the items that PR deliberately left out, and the first one is load-bearing for two others.
The batch await makes concurrency accounting dead code
poll() fully awaits Promise.allSettled over each claimed batch, and schedulePoll only runs after the poll resolves. So activeJobs.size and activeJobsByType are always 0 when availableSlots and typeLimit - activeForType are computed. Both expressions are unreachable: mutating each to ignore in-flight work passes the whole suite.
It does not over-fetch — the serialization accidentally bounds it — but per-type concurrency limits never bind, which is not what concurrencyByType looks like it does.
Which is also why /health is still wrong (#182, #184)
Because poll() blocks for the whole batch duration, lastPollTime stalls for as long as the longest job. buildHealthResponse compares it against STALE_POLL_MS = 60_000 with no active-job exemption, against a 120s AI_RESPONSE timeout. So a healthy worker mid-job returns 503 "stalled" and Railway restarts it — killing the job it was in the middle of.
Two ways out, and they should be decided together:
- Stop awaiting the batch (fixes the accounting too), or
- Exempt active jobs from the stale-poll check.
(1) is the better fix and makes (2) unnecessary. This is the reason /health has now been deferred three times; it should be one PR with the batch change.
An unregistered job type is terminally destroyed
FAILED with completedAt and no attempt increment — and job-cleanup.ts only deletes COMPLETED and DEAD_LETTER, so the rows accumulate forever. #224 added lockedAt: null, claimToken: null to that write, which means reclaim can't rescue them either.
Reachable only via claimAndProcessJobs, and production sets concurrencyByType so it takes the type-filtered path. Latent today; live the moment concurrencyByType is emptied or a worker ships without the full handler set.
A no-op release back to PENDING with no attempt consumed is the safe failure. Data-destroying path, so it wants its own change.
The timeout doesn't cancel the handler, and the drain has no deadline
runWithTimeout stops waiting for the handler; the handler keeps running. And AI_RESPONSE at 120s against Railway's ~30s SIGTERM grace means the graceful drain resolves well after SIGKILL already landed.
So the real shutdown path is still crash-abandonment plus reclaim. #224's fence makes that safe at the row level — it is not cancellation, and shouldn't be read as such. Threading an AbortSignal through JobHandlerContext is the actual fix.
Handlers with external side effects need idempotency keys
A database fence can't un-send an email. In an interleaving where a live claim is reclaimed, the first execution's external effects have already happened. lockUntil (#224) makes that much rarer, not impossible.
AI_RESPONSE— covered by #191'sMessage.ticketId_responseKeyand theresponseStatemachine.ESCALATION,HUBSPOT_SYNC,TRACKER_SYNC— not covered, and #224 is what introduces automatic re-execution.
Smaller
void currentPoll.finally(cb)builds a derived promise with no rejection handler.poll()can't reject today, so latent — but if a throw ever escapes it's an unhandled rejection, andstop()'sawait this.pollPromisewould propagate out so$disconnect()never runs.job-cleanup.tsnever deletesFAILEDrows at all, which is what makes the accumulation above unbounded.
- 主要言語
- TypeScript
- スター
- 8
- フォーク
- 4
- 平均マージ
- 5日 4時間
- マージ済み PR(30日)
- 11
環境構築
- Dockerfile または Docker Compose ファイルあり
- プルリクエストのテンプレートなし
- コントリビューションガイドなし
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
CopilotKit/outpost のほかの issue
-
area: docs area: security roadmap: now
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
CopilotKit/outpost#277 ·
メンテナーはふだん 7 日以内に返信
-
area: infrastructure roadmap roadmap: later
難易度 1/5 1時間未満 初心者へのやさしさ 74/100
CopilotKit/outpost#179 ·
メンテナーはふだん 7 日以内に返信
-
area: ai roadmap roadmap: now
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
CopilotKit/outpost#145 ·
メンテナーはふだん 7 日以内に返信
-
area: integrations priority: low roadmap roadmap: later
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
CopilotKit/outpost#124 · コメント 3 件 ·
メンテナーはふだん 7 日以内に返信
-
area: integrations priority: low roadmap roadmap: later
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
CopilotKit/outpost#123 · コメント 2 件 ·
メンテナーはふだん 7 日以内に返信
CopilotKit/outpost の issue をすべて見る
似ている issue
-
bug
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
StabilityNexus/Fate-EVM-Frontend#153 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
code-yeongyu/oh-my-openagent#9039 ·
メンテナーはふだん 1 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
Tencent/teamai-cli#862 ·
メンテナーはふだん 1 日以内に返信
-
bug good first issue hacktoberfest redis
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
libredb/libredb-studio#1164 ·
メンテナーはふだん 1 日以内に返信
-
flake
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
メンテナーはふだん 1 日以内に返信