Queue keeper: hold only rows due within a configurable window (queue.keeper.window)
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 38/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Stack tecnologico
- javascript
- Ambito
- backend
Direzione di ricerca
Start with apply, topK, dueSummary, and state in packages/plugin/src/util/queueKeeper.js, then trace load, verify, publish, and the touched protocol in queueKeeperService.js. Review deriveQueueStatus and claimSchedules in renderSchedule.js, plus configSchema.js and backlogSnapshot.js, before running the listed queue keeper and service tests. Done means the configurable window, edge and verification behavior work together, with correct status, counts, metrics, and tests for partial or failed walks.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Today the queue keeper holds every schedule row its node owns in memory. This issue proposes an option to hold only the rows due within a configurable window ahead of now. Rows due later stay in RenderSchedule only, and the keeper takes each one in before it comes due.
TL;DR
- Add
queue.keeper.window(ms). The default,0, holds everything, as today. Turning it on is a config change for each deployment. - The catch: since #219 removed the
nextRenderTimeindex, a row the keeper does not hold can only come back in two ways. Something writes it (an event), or the full primary-key walk reads it. With a window, the verification walk is no longer just a safety net: for rows nobody writes, it is the only way into the window. - So the keeper tracks an edge: every owned row due before it is held. It must report "not complete" once the clock passes the edge, never
emptywhile work it cannot see is due. - No count cap (
maxRows) in this issue. A cap set above the due set adds little over a time window. A cap that cuts into the due set breaks claim priority and the exactduecount. See "Why not a count cap" below.
Why
| Rows held | Worker 0 heap | |
|---|---|---|
| A production deployment today (about 470k rows per node, observed at the #219 rollout) | all | ~94 MB |
| Same, with a 24h window | ~27% plus the due backlog | ~25 MB plus the backlog |
| Same, with a 6h window | ~7% plus the due backlog | ~6 MB plus the backlog |
| A 10M-row node today | all | ~2 GB |
These are estimates: rows × the benchmarked ~200 B/row (#215). The held fractions assume each row's due time is spread evenly across its cadence (hypothesis). That deployment has ~97% of its rows on a 96h cadence and ~3% on 24h. A demand ladder that promotes rows to faster rungs raises the fraction.
Memory is not a problem at today's scale. A window is what makes a corpus of millions of rows per node feasible.
What does not change:
- The load and the verification walk still read every row, so walk time and I/O stay the same.
- Every write is still delivered by the subscription. An event beyond the window can skip
classify, becauseapplyalready computes the due minute first (util/queueKeeper.js:220-221).
Design
- One admission rule. Hold a row if it is due before
now + windowat the moment it is applied. The rule is the same for walked rows and for events. A held row leaves only when an event moves it out of the window or deletes it. - The edge. Say a walk starts at t0 and finishes. Every owned row due before
t0 + windowis then held, so the keeper recordsedge = t0 + window.- Each row was either walked after t0, and so admitted by the rule, or touched by an event after t0. The existing
touchedprotocol already makes the walk skip a touched row in favour of its event. - The window only moves forward, and only an event can push a row past the edge.
- A partial walk (rows it could not read) still advances the edge: it read every readable row. A partial load today already serves without the unreadable rows, and
exactstays false as it does now.
- Each row was either walked after t0, and so admitted by the rule, or touched by an event after t0. The existing
- Keep the edge ahead of the clock.
- With a window set, the walk runs every
min(verifyInterval, window / 3). verifyInterval: 0no longer disables the walk when a window is set: it is load-bearing.- A failed walk retries early, through the existing
walkRetrypath.
- With a window set, the walk runs every
- When the clock passes the edge. Rows nobody wrote that come due after the edge are invisible: they are never published and never claimed. Today
deriveQueueStatusreportsemptywhenever the held due count is 0 andcompleteis set (util/renderSchedule.js:592-597). So:complete(in the keeper signal and inclaimSchedules) becomes "the load finished and now < edge".- Past the edge, the node reports
queuedwhile it still holds due rows, andunreadyonce it doesn't, neverempty. exactturns false.- Log an error, and emit a
queue_healthseries for how far ahead of now the edge is.
- The verification counts. Rows the walk takes into the window are admissions, not repairs. Only an owned row due before the previous edge and not held counts as
missing. Otherwisekeeper_repairedis non-zero on every walk andexactnever comes back. That is the same trap as the unschedulable-row fix atutil/queueKeeperService.js:707. Count admissions as their own field. This sits next to the #215 checklist item that splits rows gained after a membership change from missed writes. - Unchanged paths. Reclassify, resync and membership changes work on held rows. A row outside the window is classified when it is admitted, and the verification walk after a membership change admits the newly owned rows inside the window.
What the counts mean with a window
| Queue-state field | With a window |
|---|---|
due, lateness, byRoute, classes, oldest*, changedByDemand |
exact while now < edge (every due row is held) |
coming.next15m, coming.next60m |
exact while the window is at least 60 min |
coming.next24h, coming.byHour (24 hourly bins) |
cut off at the window. These feed upcomingFromKeeper (util/backlogSnapshot.js:56) and the console's queue view |
rows, keeper.rows |
rows held, not rows owned |
Fix for the last two rows: the walk counts the owned rows it does not hold, per hour out to 24h, plus an owned total. Report them stamped with that walk's time, as approximate. The alternative is a window of at least 24h.
Why not a count cap
| Cap position | Effect |
|---|---|
| Above the due set | Works the same as a window whose edge is the Nth row's due minute. Adds little over window. |
| Into the due set | It holds the earliest-due rows, but claim order is lateness ÷ cadence (util/renderPriority.js), before the sitemap boost. A 96h row two days late scores 0.5. A 1h row two hours late scores 2.0, and it may be the row not held: priority inverts until the next walk. due, which autoscaling reads, becomes a lower bound. Once the held due rows drain while the edge is behind now, the node sits idle with work due unless an extra refill walk runs. |
The due backlog is the part a window cannot bound: a fleet outage or a bulk re-file. Revisit a cap only if a real corpus needs a bound there, and then as an explicitly inexact mode.
Tests
test/queueKeeper.test.js: the admission rule, a row dropped when an event moves it out of the window, andwindow: 0holding everything.- Service tests:
- a finished walk advances the edge; a partial walk advances it too, with
exactfalse - a failed walk lets the clock pass the edge:
completeturns false, and the status isunready, notempty - admissions are not counted as
missing - the walk interval is clamped to window / 3
- a finished walk advances the edge; a partial walk advances it too, with
Open items
- untested: walk time at millions of rows per node. Observed: the keeper went live 13–17 s after a restart at ~470k rows per node in production. With a window, that walk must finish well inside
window / 3. - Decide: the default.
0(off) is proposed. A window of 24h or more keeps every count exact without the counts from step 5 of the design. - Decide: whether an event beyond the window should still be counted in the flow history. Today it counts as
removedwhen it drops a held row. A row that was never held counts nothing. - Version: claim the version when work starts. Open PRs have taken 0.97.3 (#230), 0.97.4 (#231) and 0.98.0 (#232). None of the three touches the keeper files. #232 touches
configSchema.js(route options), so a conflict there is unlikely.
Sources
packages/plugin/src/util/queueKeeper.js:apply,topK,dueSummary,statepackages/plugin/src/util/queueKeeperService.js: load,verify,publish, thetouchedprotocolpackages/plugin/src/util/renderSchedule.js: the keeper signal,deriveQueueStatus,claimSchedulespackages/plugin/src/util/backlogSnapshot.js:upcomingFromKeeperpackages/plugin/src/configSchema.js:queue.keeper(verifyIntervaldefaults to 60 min, and0disables it)- #215 (keeper design and benchmarks), #219 (built, prerender-v0.93.0)
🤖 Generated with Claude Code
- Lingua principale
- JavaScript
- Stelle
- 0
- Fork
- 0
- Merge medio
- 8h 11m
- PR unite (30g)
- 73
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Nessun modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di HarperFast/prerender-plugin
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
HarperFast/prerender-plugin#245 ·
I maintainer di solito rispondono entro 1 giorno
-
enhancement
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
HarperFast/prerender-plugin#244 ·
I maintainer di solito rispondono entro 1 giorno
-
enhancement
Difficoltà 5/5 Più di una settimana Idoneità per principianti 30/100
HarperFast/prerender-plugin#242 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
enhancement performance
Difficoltà 5/5 Più di una settimana Idoneità per principianti 38/100
HarperFast/prerender-plugin#235 ·
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 3/5 1-2 giorni Idoneità per principianti 74/100
HarperFast/prerender-plugin#218 ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di HarperFast/prerender-plugin
Issue simili
-
external-issue to-triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
LearningCircuit/local-deep-research#7067 ·
I maintainer di solito rispondono entro 1 giorno
-
automated issue report
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 65/100
lirantal/discoprint#32 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
I maintainer di solito rispondono entro 1 giorno