Queue keeper: hold only rows due within a configurable window (queue.keeper.window)
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 5/5
- 预计耗时
- 一周以上
- 新手友好度
- 38/100
- Issue 类型
- 功能
- 描述清晰度
- 基本清楚
- 活跃度
- 活跃
- 技术栈
- javascript
- 领域
- backend
调研方向
Start with apply, topK, dueSummary, and state in packages/plugin/src/util/queueKeeper.js, then trace load, verify, publish, and the touched protocol in queueKeeperService.js. Review deriveQueueStatus and claimSchedules in renderSchedule.js, plus configSchema.js and backlogSnapshot.js, before running the listed queue keeper and service tests. Done means the configurable window, edge and verification behavior work together, with correct status, counts, metrics, and tests for partial or failed walks.
由索引模型根据 Issue 内容生成。
描述
Today the queue keeper holds every schedule row its node owns in memory. This issue proposes an option to hold only the rows due within a configurable window ahead of now. Rows due later stay in RenderSchedule only, and the keeper takes each one in before it comes due.
TL;DR
- Add
queue.keeper.window(ms). The default,0, holds everything, as today. Turning it on is a config change for each deployment. - The catch: since #219 removed the
nextRenderTimeindex, a row the keeper does not hold can only come back in two ways. Something writes it (an event), or the full primary-key walk reads it. With a window, the verification walk is no longer just a safety net: for rows nobody writes, it is the only way into the window. - So the keeper tracks an edge: every owned row due before it is held. It must report "not complete" once the clock passes the edge, never
emptywhile work it cannot see is due. - No count cap (
maxRows) in this issue. A cap set above the due set adds little over a time window. A cap that cuts into the due set breaks claim priority and the exactduecount. See "Why not a count cap" below.
Why
| Rows held | Worker 0 heap | |
|---|---|---|
| A production deployment today (about 470k rows per node, observed at the #219 rollout) | all | ~94 MB |
| Same, with a 24h window | ~27% plus the due backlog | ~25 MB plus the backlog |
| Same, with a 6h window | ~7% plus the due backlog | ~6 MB plus the backlog |
| A 10M-row node today | all | ~2 GB |
These are estimates: rows × the benchmarked ~200 B/row (#215). The held fractions assume each row's due time is spread evenly across its cadence (hypothesis). That deployment has ~97% of its rows on a 96h cadence and ~3% on 24h. A demand ladder that promotes rows to faster rungs raises the fraction.
Memory is not a problem at today's scale. A window is what makes a corpus of millions of rows per node feasible.
What does not change:
- The load and the verification walk still read every row, so walk time and I/O stay the same.
- Every write is still delivered by the subscription. An event beyond the window can skip
classify, becauseapplyalready computes the due minute first (util/queueKeeper.js:220-221).
Design
- One admission rule. Hold a row if it is due before
now + windowat the moment it is applied. The rule is the same for walked rows and for events. A held row leaves only when an event moves it out of the window or deletes it. - The edge. Say a walk starts at t0 and finishes. Every owned row due before
t0 + windowis then held, so the keeper recordsedge = t0 + window.- Each row was either walked after t0, and so admitted by the rule, or touched by an event after t0. The existing
touchedprotocol already makes the walk skip a touched row in favour of its event. - The window only moves forward, and only an event can push a row past the edge.
- A partial walk (rows it could not read) still advances the edge: it read every readable row. A partial load today already serves without the unreadable rows, and
exactstays false as it does now.
- Each row was either walked after t0, and so admitted by the rule, or touched by an event after t0. The existing
- Keep the edge ahead of the clock.
- With a window set, the walk runs every
min(verifyInterval, window / 3). verifyInterval: 0no longer disables the walk when a window is set: it is load-bearing.- A failed walk retries early, through the existing
walkRetrypath.
- With a window set, the walk runs every
- When the clock passes the edge. Rows nobody wrote that come due after the edge are invisible: they are never published and never claimed. Today
deriveQueueStatusreportsemptywhenever the held due count is 0 andcompleteis set (util/renderSchedule.js:592-597). So:complete(in the keeper signal and inclaimSchedules) becomes "the load finished and now < edge".- Past the edge, the node reports
queuedwhile it still holds due rows, andunreadyonce it doesn't, neverempty. exactturns false.- Log an error, and emit a
queue_healthseries for how far ahead of now the edge is.
- The verification counts. Rows the walk takes into the window are admissions, not repairs. Only an owned row due before the previous edge and not held counts as
missing. Otherwisekeeper_repairedis non-zero on every walk andexactnever comes back. That is the same trap as the unschedulable-row fix atutil/queueKeeperService.js:707. Count admissions as their own field. This sits next to the #215 checklist item that splits rows gained after a membership change from missed writes. - Unchanged paths. Reclassify, resync and membership changes work on held rows. A row outside the window is classified when it is admitted, and the verification walk after a membership change admits the newly owned rows inside the window.
What the counts mean with a window
| Queue-state field | With a window |
|---|---|
due, lateness, byRoute, classes, oldest*, changedByDemand |
exact while now < edge (every due row is held) |
coming.next15m, coming.next60m |
exact while the window is at least 60 min |
coming.next24h, coming.byHour (24 hourly bins) |
cut off at the window. These feed upcomingFromKeeper (util/backlogSnapshot.js:56) and the console's queue view |
rows, keeper.rows |
rows held, not rows owned |
Fix for the last two rows: the walk counts the owned rows it does not hold, per hour out to 24h, plus an owned total. Report them stamped with that walk's time, as approximate. The alternative is a window of at least 24h.
Why not a count cap
| Cap position | Effect |
|---|---|
| Above the due set | Works the same as a window whose edge is the Nth row's due minute. Adds little over window. |
| Into the due set | It holds the earliest-due rows, but claim order is lateness ÷ cadence (util/renderPriority.js), before the sitemap boost. A 96h row two days late scores 0.5. A 1h row two hours late scores 2.0, and it may be the row not held: priority inverts until the next walk. due, which autoscaling reads, becomes a lower bound. Once the held due rows drain while the edge is behind now, the node sits idle with work due unless an extra refill walk runs. |
The due backlog is the part a window cannot bound: a fleet outage or a bulk re-file. Revisit a cap only if a real corpus needs a bound there, and then as an explicitly inexact mode.
Tests
test/queueKeeper.test.js: the admission rule, a row dropped when an event moves it out of the window, andwindow: 0holding everything.- Service tests:
- a finished walk advances the edge; a partial walk advances it too, with
exactfalse - a failed walk lets the clock pass the edge:
completeturns false, and the status isunready, notempty - admissions are not counted as
missing - the walk interval is clamped to window / 3
- a finished walk advances the edge; a partial walk advances it too, with
Open items
- untested: walk time at millions of rows per node. Observed: the keeper went live 13–17 s after a restart at ~470k rows per node in production. With a window, that walk must finish well inside
window / 3. - Decide: the default.
0(off) is proposed. A window of 24h or more keeps every count exact without the counts from step 5 of the design. - Decide: whether an event beyond the window should still be counted in the flow history. Today it counts as
removedwhen it drops a held row. A row that was never held counts nothing. - Version: claim the version when work starts. Open PRs have taken 0.97.3 (#230), 0.97.4 (#231) and 0.98.0 (#232). None of the three touches the keeper files. #232 touches
configSchema.js(route options), so a conflict there is unlikely.
Sources
packages/plugin/src/util/queueKeeper.js:apply,topK,dueSummary,statepackages/plugin/src/util/queueKeeperService.js: load,verify,publish, thetouchedprotocolpackages/plugin/src/util/renderSchedule.js: the keeper signal,deriveQueueStatus,claimSchedulespackages/plugin/src/util/backlogSnapshot.js:upcomingFromKeeperpackages/plugin/src/configSchema.js:queue.keeper(verifyIntervaldefaults to 60 min, and0disables it)- #215 (keeper design and benchmarks), #219 (built, prerender-v0.93.0)
🤖 Generated with Claude Code
- 主要语言
- JavaScript
- 星标
- 0
- 派生
- 0
- 平均合并
- 8 小时 25 分钟
- 30 天内合并 PR
- 71
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
HarperFast/prerender-plugin 的其他 Issue
-
enhancement
难度 5/5 一周以上 新手友好度 25/100
HarperFast/prerender-plugin#244 ·
维护者通常 1 天内回复
-
enhancement
难度 5/5 一周以上 新手友好度 30/100
HarperFast/prerender-plugin#242 · 1 条评论 ·
维护者通常 1 天内回复
-
enhancement performance
难度 5/5 一周以上 新手友好度 38/100
HarperFast/prerender-plugin#235 ·
维护者通常 1 天内回复
-
bug
难度 3/5 1-2 天 新手友好度 74/100
HarperFast/prerender-plugin#218 ·
维护者通常 1 天内回复
-
难度 5/5 一周以上 新手友好度 15/100
HarperFast/prerender-plugin#215 ·
维护者通常 1 天内回复
查看 HarperFast/prerender-plugin 的全部 Issue
相似的 Issue
-
[quality] useFocusTrap's Shift+Tab wrap and non-Tab/non-Escape key arms are never driven end to end可能已有人在做 @hivecommons-hive 今天认领。 未关闭agent/quality hive/covered-by-pr hive/hosted-available-lke648397-260827-5n31 quality testing
难度 2/5 1-3 小时 新手友好度 85/100
维护者通常 1 天内回复
-
agentic-workflows
难度 1/5 1 小时以内 新手友好度 85/100
githubnext/gh-aw-workshop#4220 ·
维护者通常 1 天内回复
-
难度 1/5 1 小时以内 新手友好度 92/100
JuliusBrussee/caveman#1189 ·
维护者通常 1 天内回复
-
priority:low ready-for-dev
难度 2/5 1-3 小时 新手友好度 78/100
OpenHands/extensions#738 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 82/100
invoiceninja/invoiceninja#13320 ·
维护者通常 1 天内回复