Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Queue keeper: hold only rows due within a configurable window (queue.keeper.window)

Aperta
#233 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
5/5
Tempo stimato
Più di una settimana
Idoneità per principianti
38/100
Tipo di issue
Funzionalità
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
javascript
Ambito
backend

Direzione di ricerca

Start with apply, topK, dueSummary, and state in packages/plugin/src/util/queueKeeper.js, then trace load, verify, publish, and the touched protocol in queueKeeperService.js. Review deriveQueueStatus and claimSchedules in renderSchedule.js, plus configSchema.js and backlogSnapshot.js, before running the listed queue keeper and service tests. Done means the configurable window, edge and verification behavior work together, with correct status, counts, metrics, and tests for partial or failed walks.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

enhancement performance

Today the queue keeper holds every schedule row its node owns in memory. This issue proposes an option to hold only the rows due within a configurable window ahead of now. Rows due later stay in RenderSchedule only, and the keeper takes each one in before it comes due.

TL;DR

  • Add queue.keeper.window (ms). The default, 0, holds everything, as today. Turning it on is a config change for each deployment.
  • The catch: since #219 removed the nextRenderTime index, a row the keeper does not hold can only come back in two ways. Something writes it (an event), or the full primary-key walk reads it. With a window, the verification walk is no longer just a safety net: for rows nobody writes, it is the only way into the window.
  • So the keeper tracks an edge: every owned row due before it is held. It must report "not complete" once the clock passes the edge, never empty while work it cannot see is due.
  • No count cap (maxRows) in this issue. A cap set above the due set adds little over a time window. A cap that cuts into the due set breaks claim priority and the exact due count. See "Why not a count cap" below.

Why

Rows held Worker 0 heap
A production deployment today (about 470k rows per node, observed at the #219 rollout) all ~94 MB
Same, with a 24h window ~27% plus the due backlog ~25 MB plus the backlog
Same, with a 6h window ~7% plus the due backlog ~6 MB plus the backlog
A 10M-row node today all ~2 GB

These are estimates: rows × the benchmarked ~200 B/row (#215). The held fractions assume each row's due time is spread evenly across its cadence (hypothesis). That deployment has ~97% of its rows on a 96h cadence and ~3% on 24h. A demand ladder that promotes rows to faster rungs raises the fraction.

Memory is not a problem at today's scale. A window is what makes a corpus of millions of rows per node feasible.

What does not change:

  • The load and the verification walk still read every row, so walk time and I/O stay the same.
  • Every write is still delivered by the subscription. An event beyond the window can skip classify, because apply already computes the due minute first (util/queueKeeper.js:220-221).

Design

  1. One admission rule. Hold a row if it is due before now + window at the moment it is applied. The rule is the same for walked rows and for events. A held row leaves only when an event moves it out of the window or deletes it.
  2. The edge. Say a walk starts at t0 and finishes. Every owned row due before t0 + window is then held, so the keeper records edge = t0 + window.
    • Each row was either walked after t0, and so admitted by the rule, or touched by an event after t0. The existing touched protocol already makes the walk skip a touched row in favour of its event.
    • The window only moves forward, and only an event can push a row past the edge.
    • A partial walk (rows it could not read) still advances the edge: it read every readable row. A partial load today already serves without the unreadable rows, and exact stays false as it does now.
  3. Keep the edge ahead of the clock.
    • With a window set, the walk runs every min(verifyInterval, window / 3).
    • verifyInterval: 0 no longer disables the walk when a window is set: it is load-bearing.
    • A failed walk retries early, through the existing walkRetry path.
  4. When the clock passes the edge. Rows nobody wrote that come due after the edge are invisible: they are never published and never claimed. Today deriveQueueStatus reports empty whenever the held due count is 0 and complete is set (util/renderSchedule.js:592-597). So:
    • complete (in the keeper signal and in claimSchedules) becomes "the load finished and now < edge".
    • Past the edge, the node reports queued while it still holds due rows, and unready once it doesn't, never empty.
    • exact turns false.
    • Log an error, and emit a queue_health series for how far ahead of now the edge is.
  5. The verification counts. Rows the walk takes into the window are admissions, not repairs. Only an owned row due before the previous edge and not held counts as missing. Otherwise keeper_repaired is non-zero on every walk and exact never comes back. That is the same trap as the unschedulable-row fix at util/queueKeeperService.js:707. Count admissions as their own field. This sits next to the #215 checklist item that splits rows gained after a membership change from missed writes.
  6. Unchanged paths. Reclassify, resync and membership changes work on held rows. A row outside the window is classified when it is admitted, and the verification walk after a membership change admits the newly owned rows inside the window.

What the counts mean with a window

Queue-state field With a window
due, lateness, byRoute, classes, oldest*, changedByDemand exact while now < edge (every due row is held)
coming.next15m, coming.next60m exact while the window is at least 60 min
coming.next24h, coming.byHour (24 hourly bins) cut off at the window. These feed upcomingFromKeeper (util/backlogSnapshot.js:56) and the console's queue view
rows, keeper.rows rows held, not rows owned

Fix for the last two rows: the walk counts the owned rows it does not hold, per hour out to 24h, plus an owned total. Report them stamped with that walk's time, as approximate. The alternative is a window of at least 24h.

Why not a count cap

Cap position Effect
Above the due set Works the same as a window whose edge is the Nth row's due minute. Adds little over window.
Into the due set It holds the earliest-due rows, but claim order is lateness ÷ cadence (util/renderPriority.js), before the sitemap boost. A 96h row two days late scores 0.5. A 1h row two hours late scores 2.0, and it may be the row not held: priority inverts until the next walk. due, which autoscaling reads, becomes a lower bound. Once the held due rows drain while the edge is behind now, the node sits idle with work due unless an extra refill walk runs.

The due backlog is the part a window cannot bound: a fleet outage or a bulk re-file. Revisit a cap only if a real corpus needs a bound there, and then as an explicitly inexact mode.

Tests

  • test/queueKeeper.test.js: the admission rule, a row dropped when an event moves it out of the window, and window: 0 holding everything.
  • Service tests:
    • a finished walk advances the edge; a partial walk advances it too, with exact false
    • a failed walk lets the clock pass the edge: complete turns false, and the status is unready, not empty
    • admissions are not counted as missing
    • the walk interval is clamped to window / 3

Open items

  • untested: walk time at millions of rows per node. Observed: the keeper went live 13–17 s after a restart at ~470k rows per node in production. With a window, that walk must finish well inside window / 3.
  • Decide: the default. 0 (off) is proposed. A window of 24h or more keeps every count exact without the counts from step 5 of the design.
  • Decide: whether an event beyond the window should still be counted in the flow history. Today it counts as removed when it drops a held row. A row that was never held counts nothing.
  • Version: claim the version when work starts. Open PRs have taken 0.97.3 (#230), 0.97.4 (#231) and 0.98.0 (#232). None of the three touches the keeper files. #232 touches configSchema.js (route options), so a conflict there is unlikely.

Sources

  • packages/plugin/src/util/queueKeeper.js: apply, topK, dueSummary, state
  • packages/plugin/src/util/queueKeeperService.js: load, verify, publish, the touched protocol
  • packages/plugin/src/util/renderSchedule.js: the keeper signal, deriveQueueStatus, claimSchedules
  • packages/plugin/src/util/backlogSnapshot.js: upcomingFromKeeper
  • packages/plugin/src/configSchema.js: queue.keeper (verifyInterval defaults to 60 min, and 0 disables it)
  • #215 (keeper design and benchmarks), #219 (built, prerender-v0.93.0)

🤖 Generated with Claude Code

Lingua principale
JavaScript
Stelle
0
Fork
0
Merge medio
8h 11m
PR unite (30g)
73

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di HarperFast/prerender-plugin

Tutte le issue di HarperFast/prerender-plugin

Issue simili

Altre issue su JavaScript

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.