Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Queue keeper: hold only rows due within a configurable window (queue.keeper.window)

未关闭
#233 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
5/5
预计耗时
一周以上
新手友好度
38/100
Issue 类型
功能
描述清晰度
基本清楚
活跃度
活跃
技术栈
javascript
领域
backend

调研方向

Start with apply, topK, dueSummary, and state in packages/plugin/src/util/queueKeeper.js, then trace load, verify, publish, and the touched protocol in queueKeeperService.js. Review deriveQueueStatus and claimSchedules in renderSchedule.js, plus configSchema.js and backlogSnapshot.js, before running the listed queue keeper and service tests. Done means the configurable window, edge and verification behavior work together, with correct status, counts, metrics, and tests for partial or failed walks.

由索引模型根据 Issue 内容生成。

描述

enhancement performance

Today the queue keeper holds every schedule row its node owns in memory. This issue proposes an option to hold only the rows due within a configurable window ahead of now. Rows due later stay in RenderSchedule only, and the keeper takes each one in before it comes due.

TL;DR

  • Add queue.keeper.window (ms). The default, 0, holds everything, as today. Turning it on is a config change for each deployment.
  • The catch: since #219 removed the nextRenderTime index, a row the keeper does not hold can only come back in two ways. Something writes it (an event), or the full primary-key walk reads it. With a window, the verification walk is no longer just a safety net: for rows nobody writes, it is the only way into the window.
  • So the keeper tracks an edge: every owned row due before it is held. It must report "not complete" once the clock passes the edge, never empty while work it cannot see is due.
  • No count cap (maxRows) in this issue. A cap set above the due set adds little over a time window. A cap that cuts into the due set breaks claim priority and the exact due count. See "Why not a count cap" below.

Why

Rows held Worker 0 heap
A production deployment today (about 470k rows per node, observed at the #219 rollout) all ~94 MB
Same, with a 24h window ~27% plus the due backlog ~25 MB plus the backlog
Same, with a 6h window ~7% plus the due backlog ~6 MB plus the backlog
A 10M-row node today all ~2 GB

These are estimates: rows × the benchmarked ~200 B/row (#215). The held fractions assume each row's due time is spread evenly across its cadence (hypothesis). That deployment has ~97% of its rows on a 96h cadence and ~3% on 24h. A demand ladder that promotes rows to faster rungs raises the fraction.

Memory is not a problem at today's scale. A window is what makes a corpus of millions of rows per node feasible.

What does not change:

  • The load and the verification walk still read every row, so walk time and I/O stay the same.
  • Every write is still delivered by the subscription. An event beyond the window can skip classify, because apply already computes the due minute first (util/queueKeeper.js:220-221).

Design

  1. One admission rule. Hold a row if it is due before now + window at the moment it is applied. The rule is the same for walked rows and for events. A held row leaves only when an event moves it out of the window or deletes it.
  2. The edge. Say a walk starts at t0 and finishes. Every owned row due before t0 + window is then held, so the keeper records edge = t0 + window.
    • Each row was either walked after t0, and so admitted by the rule, or touched by an event after t0. The existing touched protocol already makes the walk skip a touched row in favour of its event.
    • The window only moves forward, and only an event can push a row past the edge.
    • A partial walk (rows it could not read) still advances the edge: it read every readable row. A partial load today already serves without the unreadable rows, and exact stays false as it does now.
  3. Keep the edge ahead of the clock.
    • With a window set, the walk runs every min(verifyInterval, window / 3).
    • verifyInterval: 0 no longer disables the walk when a window is set: it is load-bearing.
    • A failed walk retries early, through the existing walkRetry path.
  4. When the clock passes the edge. Rows nobody wrote that come due after the edge are invisible: they are never published and never claimed. Today deriveQueueStatus reports empty whenever the held due count is 0 and complete is set (util/renderSchedule.js:592-597). So:
    • complete (in the keeper signal and in claimSchedules) becomes "the load finished and now < edge".
    • Past the edge, the node reports queued while it still holds due rows, and unready once it doesn't, never empty.
    • exact turns false.
    • Log an error, and emit a queue_health series for how far ahead of now the edge is.
  5. The verification counts. Rows the walk takes into the window are admissions, not repairs. Only an owned row due before the previous edge and not held counts as missing. Otherwise keeper_repaired is non-zero on every walk and exact never comes back. That is the same trap as the unschedulable-row fix at util/queueKeeperService.js:707. Count admissions as their own field. This sits next to the #215 checklist item that splits rows gained after a membership change from missed writes.
  6. Unchanged paths. Reclassify, resync and membership changes work on held rows. A row outside the window is classified when it is admitted, and the verification walk after a membership change admits the newly owned rows inside the window.

What the counts mean with a window

Queue-state field With a window
due, lateness, byRoute, classes, oldest*, changedByDemand exact while now < edge (every due row is held)
coming.next15m, coming.next60m exact while the window is at least 60 min
coming.next24h, coming.byHour (24 hourly bins) cut off at the window. These feed upcomingFromKeeper (util/backlogSnapshot.js:56) and the console's queue view
rows, keeper.rows rows held, not rows owned

Fix for the last two rows: the walk counts the owned rows it does not hold, per hour out to 24h, plus an owned total. Report them stamped with that walk's time, as approximate. The alternative is a window of at least 24h.

Why not a count cap

Cap position Effect
Above the due set Works the same as a window whose edge is the Nth row's due minute. Adds little over window.
Into the due set It holds the earliest-due rows, but claim order is lateness ÷ cadence (util/renderPriority.js), before the sitemap boost. A 96h row two days late scores 0.5. A 1h row two hours late scores 2.0, and it may be the row not held: priority inverts until the next walk. due, which autoscaling reads, becomes a lower bound. Once the held due rows drain while the edge is behind now, the node sits idle with work due unless an extra refill walk runs.

The due backlog is the part a window cannot bound: a fleet outage or a bulk re-file. Revisit a cap only if a real corpus needs a bound there, and then as an explicitly inexact mode.

Tests

  • test/queueKeeper.test.js: the admission rule, a row dropped when an event moves it out of the window, and window: 0 holding everything.
  • Service tests:
    • a finished walk advances the edge; a partial walk advances it too, with exact false
    • a failed walk lets the clock pass the edge: complete turns false, and the status is unready, not empty
    • admissions are not counted as missing
    • the walk interval is clamped to window / 3

Open items

  • untested: walk time at millions of rows per node. Observed: the keeper went live 13–17 s after a restart at ~470k rows per node in production. With a window, that walk must finish well inside window / 3.
  • Decide: the default. 0 (off) is proposed. A window of 24h or more keeps every count exact without the counts from step 5 of the design.
  • Decide: whether an event beyond the window should still be counted in the flow history. Today it counts as removed when it drops a held row. A row that was never held counts nothing.
  • Version: claim the version when work starts. Open PRs have taken 0.97.3 (#230), 0.97.4 (#231) and 0.98.0 (#232). None of the three touches the keeper files. #232 touches configSchema.js (route options), so a conflict there is unlikely.

Sources

  • packages/plugin/src/util/queueKeeper.js: apply, topK, dueSummary, state
  • packages/plugin/src/util/queueKeeperService.js: load, verify, publish, the touched protocol
  • packages/plugin/src/util/renderSchedule.js: the keeper signal, deriveQueueStatus, claimSchedules
  • packages/plugin/src/util/backlogSnapshot.js: upcomingFromKeeper
  • packages/plugin/src/configSchema.js: queue.keeper (verifyInterval defaults to 60 min, and 0 disables it)
  • #215 (keeper design and benchmarks), #219 (built, prerender-v0.93.0)

🤖 Generated with Claude Code

主要语言
JavaScript
星标
0
派生
0
平均合并
8 小时 25 分钟
30 天内合并 PR
71

环境准备

  • 没有 Dockerfile 或 Docker Compose 文件
  • 没有 Pull Request 模板
  • 阅读贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

HarperFast/prerender-plugin 的其他 Issue

查看 HarperFast/prerender-plugin 的全部 Issue

相似的 Issue

更多 JavaScript Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。