Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Queue keeper: hold only rows due within a configurable window (queue.keeper.window)

Abierto
#233 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
38/100
Tipo de issue
Nueva funcionalidad
Claridad
Bastante claro
Estado de actividad
Activo
Stack tecnológico
javascript
Área
backend

Línea de trabajo

Start with apply, topK, dueSummary, and state in packages/plugin/src/util/queueKeeper.js, then trace load, verify, publish, and the touched protocol in queueKeeperService.js. Review deriveQueueStatus and claimSchedules in renderSchedule.js, plus configSchema.js and backlogSnapshot.js, before running the listed queue keeper and service tests. Done means the configurable window, edge and verification behavior work together, with correct status, counts, metrics, and tests for partial or failed walks.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

enhancement performance

Today the queue keeper holds every schedule row its node owns in memory. This issue proposes an option to hold only the rows due within a configurable window ahead of now. Rows due later stay in RenderSchedule only, and the keeper takes each one in before it comes due.

TL;DR

  • Add queue.keeper.window (ms). The default, 0, holds everything, as today. Turning it on is a config change for each deployment.
  • The catch: since #219 removed the nextRenderTime index, a row the keeper does not hold can only come back in two ways. Something writes it (an event), or the full primary-key walk reads it. With a window, the verification walk is no longer just a safety net: for rows nobody writes, it is the only way into the window.
  • So the keeper tracks an edge: every owned row due before it is held. It must report "not complete" once the clock passes the edge, never empty while work it cannot see is due.
  • No count cap (maxRows) in this issue. A cap set above the due set adds little over a time window. A cap that cuts into the due set breaks claim priority and the exact due count. See "Why not a count cap" below.

Why

Rows held Worker 0 heap
A production deployment today (about 470k rows per node, observed at the #219 rollout) all ~94 MB
Same, with a 24h window ~27% plus the due backlog ~25 MB plus the backlog
Same, with a 6h window ~7% plus the due backlog ~6 MB plus the backlog
A 10M-row node today all ~2 GB

These are estimates: rows × the benchmarked ~200 B/row (#215). The held fractions assume each row's due time is spread evenly across its cadence (hypothesis). That deployment has ~97% of its rows on a 96h cadence and ~3% on 24h. A demand ladder that promotes rows to faster rungs raises the fraction.

Memory is not a problem at today's scale. A window is what makes a corpus of millions of rows per node feasible.

What does not change:

  • The load and the verification walk still read every row, so walk time and I/O stay the same.
  • Every write is still delivered by the subscription. An event beyond the window can skip classify, because apply already computes the due minute first (util/queueKeeper.js:220-221).

Design

  1. One admission rule. Hold a row if it is due before now + window at the moment it is applied. The rule is the same for walked rows and for events. A held row leaves only when an event moves it out of the window or deletes it.
  2. The edge. Say a walk starts at t0 and finishes. Every owned row due before t0 + window is then held, so the keeper records edge = t0 + window.
    • Each row was either walked after t0, and so admitted by the rule, or touched by an event after t0. The existing touched protocol already makes the walk skip a touched row in favour of its event.
    • The window only moves forward, and only an event can push a row past the edge.
    • A partial walk (rows it could not read) still advances the edge: it read every readable row. A partial load today already serves without the unreadable rows, and exact stays false as it does now.
  3. Keep the edge ahead of the clock.
    • With a window set, the walk runs every min(verifyInterval, window / 3).
    • verifyInterval: 0 no longer disables the walk when a window is set: it is load-bearing.
    • A failed walk retries early, through the existing walkRetry path.
  4. When the clock passes the edge. Rows nobody wrote that come due after the edge are invisible: they are never published and never claimed. Today deriveQueueStatus reports empty whenever the held due count is 0 and complete is set (util/renderSchedule.js:592-597). So:
    • complete (in the keeper signal and in claimSchedules) becomes "the load finished and now < edge".
    • Past the edge, the node reports queued while it still holds due rows, and unready once it doesn't, never empty.
    • exact turns false.
    • Log an error, and emit a queue_health series for how far ahead of now the edge is.
  5. The verification counts. Rows the walk takes into the window are admissions, not repairs. Only an owned row due before the previous edge and not held counts as missing. Otherwise keeper_repaired is non-zero on every walk and exact never comes back. That is the same trap as the unschedulable-row fix at util/queueKeeperService.js:707. Count admissions as their own field. This sits next to the #215 checklist item that splits rows gained after a membership change from missed writes.
  6. Unchanged paths. Reclassify, resync and membership changes work on held rows. A row outside the window is classified when it is admitted, and the verification walk after a membership change admits the newly owned rows inside the window.

What the counts mean with a window

Queue-state field With a window
due, lateness, byRoute, classes, oldest*, changedByDemand exact while now < edge (every due row is held)
coming.next15m, coming.next60m exact while the window is at least 60 min
coming.next24h, coming.byHour (24 hourly bins) cut off at the window. These feed upcomingFromKeeper (util/backlogSnapshot.js:56) and the console's queue view
rows, keeper.rows rows held, not rows owned

Fix for the last two rows: the walk counts the owned rows it does not hold, per hour out to 24h, plus an owned total. Report them stamped with that walk's time, as approximate. The alternative is a window of at least 24h.

Why not a count cap

Cap position Effect
Above the due set Works the same as a window whose edge is the Nth row's due minute. Adds little over window.
Into the due set It holds the earliest-due rows, but claim order is lateness ÷ cadence (util/renderPriority.js), before the sitemap boost. A 96h row two days late scores 0.5. A 1h row two hours late scores 2.0, and it may be the row not held: priority inverts until the next walk. due, which autoscaling reads, becomes a lower bound. Once the held due rows drain while the edge is behind now, the node sits idle with work due unless an extra refill walk runs.

The due backlog is the part a window cannot bound: a fleet outage or a bulk re-file. Revisit a cap only if a real corpus needs a bound there, and then as an explicitly inexact mode.

Tests

  • test/queueKeeper.test.js: the admission rule, a row dropped when an event moves it out of the window, and window: 0 holding everything.
  • Service tests:
    • a finished walk advances the edge; a partial walk advances it too, with exact false
    • a failed walk lets the clock pass the edge: complete turns false, and the status is unready, not empty
    • admissions are not counted as missing
    • the walk interval is clamped to window / 3

Open items

  • untested: walk time at millions of rows per node. Observed: the keeper went live 13–17 s after a restart at ~470k rows per node in production. With a window, that walk must finish well inside window / 3.
  • Decide: the default. 0 (off) is proposed. A window of 24h or more keeps every count exact without the counts from step 5 of the design.
  • Decide: whether an event beyond the window should still be counted in the flow history. Today it counts as removed when it drops a held row. A row that was never held counts nothing.
  • Version: claim the version when work starts. Open PRs have taken 0.97.3 (#230), 0.97.4 (#231) and 0.98.0 (#232). None of the three touches the keeper files. #232 touches configSchema.js (route options), so a conflict there is unlikely.

Sources

  • packages/plugin/src/util/queueKeeper.js: apply, topK, dueSummary, state
  • packages/plugin/src/util/queueKeeperService.js: load, verify, publish, the touched protocol
  • packages/plugin/src/util/renderSchedule.js: the keeper signal, deriveQueueStatus, claimSchedules
  • packages/plugin/src/util/backlogSnapshot.js: upcomingFromKeeper
  • packages/plugin/src/configSchema.js: queue.keeper (verifyInterval defaults to 60 min, and 0 disables it)
  • #215 (keeper design and benchmarks), #219 (built, prerender-v0.93.0)

🤖 Generated with Claude Code

Lenguaje dominante
JavaScript
Estrellas
0
Forks
0
Merge medio
8 h 11 min
PR fusionados (30 d)
73

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de HarperFast/prerender-plugin

Todos los issues de HarperFast/prerender-plugin

Issues similares

Más issues de JavaScript

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.