Hacktoberfest 2026: as issues que os mantenedores marcaram para outubro, abertas e boas para iniciantes. Ver issues do Hacktoberfest

Queue state: per-node in-memory queue keeper, the render queue itself (built in #219)

Aberta
#215 0 comentários 0 reações 0 responsáveis Ver no GitHub

Mantenedores costumam responder em até 1 dia

Ninguém assumiu esta issue ainda.

Avaliação

Dificuldade
5/5
Tempo estimado
Mais de uma semana
Facilidade para iniciantes
15/100
Tipo de issue
Funcionalidade
Clareza
Razoavelmente clara
Status de atividade
Estagnada
Stack de tecnologia
javascript
Domínio
api, backend

Direção de pesquisa

Start by reading PR #219, which contains the keeper, claims, queue-state response, and index removal described here. Review the linked harness work in #216 and #217 and the unchecked next steps; this issue is done only when the remaining production-scale, console, cluster-state, and harness work is addressed.

Escrita pelo modelo de indexação a partir do texto da issue.

Descrição

Summary

Expose the render queue's state cheaply, per node and cluster-wide, for the management console and as an autoscaling signal for the render fleet. The chosen design ("design C" below) keeps each node's queue in memory on worker 0, fed by a RenderSchedule subscription, and makes it the queue itself: RenderSchedule is its durable side and the nextRenderTime index is gone.

Status: merged in #219 and released as prerender-v0.93.0, with queue-state fixes in prerender-v0.93.1; not yet deployed. The claim floor, the ready-set sweep, the index claim scan and the index itself are removed in the same change, with no transitional fallback.

  • Single node: a subscription adds about 46 µs of CPU per schedule write, the in-memory copy stays exact, and memory is about 200 B per row (about 50 MB per node at production scale).
  • Two nodes: each node's keeper receives every replicated write and stays exact. Its cost is within round-to-round noise in process CPU, and adds 7–13 µs per write on the keeper's thread.
  • Real Harper, upgrading 0.92.0 → 0.93.0 in place (single node, 2,001 rows): the index drop needed no manual step; the first claim 2 s after the restart was granted; queue-state reported exactly the seeded due count; claims came in keeper order with no overlap; the verification walk repaired nothing. Rolling back rebuilt the index from the records, exact; on a multi-thread node it also needs a second restart once the rebuild finishes (harper-pro#681). Details in #219.
  • A base copy that carries any row of the table re-sends each receiving node's whole table to the keeper (about 1.2 s at 250k rows, fresh). The keeper absorbs it.

Not tested yet: a churned production-size store (it now decides only how long a restart serves a partial order), more than two nodes, a live membership change, and the peer check against harper-pro.

Queue state (GET /prerender_admin/queue-state, per node)

Group Fields
now due (sitemap / discovered), in flight (leases), unclaimed (estimate), paused, status
coming due in the next 15 / 60 minutes and 24 h; next 24 h by hour
lateness due rows by lateness relative to cadence, split by sitemap vs discovered and by route; oldest due row per class. A single global head age would only measure how long the lowest-priority class has starved.
flow per minute for the last hour: came due, added, triggered, rescheduled, removed
trust live, phase, exact, state age, the keeper's load / publish / verification stats

503 (with trust and the live now fields, no counts) whenever the keeper cannot vouch for its numbers. Not built yet: the cluster fan-out and the console view.

Options considered

Change Freshness Notes
A. Publish what sweepReadySet already counted Small Sweep interval (5 min) Counts from the sweep's walk, capped at sweepCap.
B. Subscription-driven exact counters beside the old queue Medium Seconds The first step of C.
C. In-memory queue per node (chosen) Large Seconds The keeper decides claim order and refills the shared-memory ready set; the sweep, the claim floor and the index go.

Design C, as built

  1. Load, serving from the first chunk. A node whose system.hdb_nodes names another node waits (up to 2 minutes) to see one, since ownership is not knowable before the node list is; a single node does not wait. Subscribe (listener form, omitCurrent: true), then walk the node's own rows by primary key in local chunks, publishing the ready set from what is held so far. Events apply as they arrive; a walked row is applied only if no event touched its key since the walk began. If an unreadable row stops the walk, the rest is read from the top down.
  2. Hold key → (due minute, class), and per class minute → Set(keys). A class is route × cadence (and whether it came off the row) × sitemap flag. Only rows this node owns by residency are held.
  3. Stay current. Apply each put / delete / invalidate event by moving the key. With residency pinned by id, a node's subscription gets a put for every write to a row it owns and a no-op delete for each write it makes to a row another node owns (measured, see Benchmark).
  4. Serve claims. Every second, write the best few thousand keys into the shared ready set. A claim, on any worker, takes entries in order, point-reads each one's row locally, skips any no longer due, and leases the rest with an exclusive grant (no mutex; see #218). A key whose leases keep expiring with no result is held back in the lease table, 2 leases then 4, 8, … capped at its cadence. Until its first publish the node grants no claims and reports the queue status unready; the keeper reports each change of status itself, so the fleet is woken by it. A stalled keeper is still served from until its last set drains.
  5. Never reload once live. A route or default-interval change reclassifies held rows in memory; a membership change reclassifies (dropping rows no longer owned) and runs the verification walk (adding rows gained); a closed or failed subscription is reopened with backoff and walked.
  6. Check itself. Each publish re-reads the top 64 published rows (each key at most once a minute); a verification walk every hour compares every row with the table. Both repair what is wrong, and a repair never overwrites a newer event.
  7. Report state from the buckets, every 5 s, into a shared buffer any worker serves.

Priority order without rescoring. Rank is lateness relative to the row's own cadence, with a sitemap boost. Two rows in the same class share cadence and boost, so the one due earlier always ranks higher: each class list is already in priority order, and the best K is a merge over class heads.

What was given up with the index: ad-hoc due-time queries against the table. A rollback to 0.92.0 re-adds the index, which Harper rebuilds from the records (observed); 0.92.0's claims fail until the rebuild finishes.

Benchmark

Harness: bench/queue-keeper (#216); two-node harness in bench/queue-keeper/cluster (#217). The README has the method and caveats. Both run Harper 5.2.13, the version production runs, on a fresh corpus, so absolute costs are floors and ratios transfer.

Single node
  • harperfast/harper:5.2.13, 250k rows, 2 worker threads (writer + keeper).
  • 50k single-row writes per arm, 5 rounds, interleaved.
No subscription Subscription Subscription + keeper
put p50 / p99 92 / 183 µs 97 / 188 µs 98 / 212 µs
process CPU per write 106 µs 142 µs 152 µs
keeper thread busy per write — 26 µs 31 µs
events per write / lag p99 — 1 / 0.8 ms 1 / 0.8 ms
  • Correctness: the keeper matched the table on 250,000 of 250,000 rows. A subscription on one thread sees every commit made on another.
  • Memory: about 136 B/row of key strings plus 64–80 B/row of structures. That's 48 MB at 250k rows and 191 MB at 1M.
  • Claim and state costs: top 5,000 due rows in 0.16–0.25 ms; full counts under 2 ms; rebuild 0.9 s for 250k rows (expect 20–30× on a churned store).
Two nodes
  • harperfast/harper-pro:5.2.13, two containers replicating over TLS, 2 worker threads each.
  • 250k rows per node, pinned the way production pins RenderSchedule (setResidencyById, rendezvous hashing with production's hash).
  • Both nodes write rows of both owners at once: 50k single-row writes per node per arm, about 15k cluster writes/s. 3 rounds, interleaved.
Per node No subscription Keeper
process CPU per cluster write 134 µs (123–141) 137 µs (133–138)
keeper thread (worker 0) busy per cluster write 44–51 µs 55–59 µs
put p50 / p99, own row 131–135 / 267 µs 126–132 / 240–247 µs
  • The owner's subscription sees every write replicated from a peer (a coverage check, 8 of 8 windows).
    • Every distinct row written to a node, by either node, reached its keeper.
    • deletes matched the node's own writes to foreign rows exactly, and no put arrived for a foreign row.
    • The 1–4 puts per 50k that never arrived were superseded versions: Harper sends only a row's current version.
  • Exact at the end: 500,000 of 500,000 rows, each on its owner, and each keeper equal to its own table.
  • Lag, writer's commit to the owner's listener: replicated p50 0.3–0.4 ms, p99 42–105 ms at 15k cluster writes/s.
  • Cost:
    • Process CPU, paired per round (keeper minus no subscription): −4 to +12 µs per cluster write, a spread as large as the effect.
    • The keeper's thread rose consistently, by 7–13 µs per cluster write, which is 10–17 µs per delivered event against 31 µs on a single node.
    • The keeper's own work was 4.5–4.8 µs per put.
  • The keeper shares its thread with replication: the table's replication socket was on worker 0 on both nodes.

What a join costs (the reload burst). add_node requested a full copy each way, and row counts never changed. A live keeper's experience depended only on whether a copy carried any row of the table (12 runs):

  • The trigger was one deleted row. Its entry crossed despite residency, and the receiver skipped storing it. Having received a row, though, the receiver wrote a reload marker, and Harper re-sent every live subscriber the node's whole table. At 250k rows that took about 1.2 s per node, and 0.4 s of keeper work.
  • With no row in the copy, nothing was re-sent.
  • A commit on the keeper's thread made no difference.
  • In production (inference): RenderSchedule rows get deleted, so expect every base copy of render_schedule to re-send each receiver's table. Base copies happen on a join, a clone, an interrupted first copy, or a peer falling behind the retained audit log.
Harper behaviour this relies on (5.2.13 source, confirmed by the runs)
  • The subscription work runs on the subscribing thread; commits are coalesced via setImmediate.
  • Harper re-reads the row before delivery, drops stale versions, and sends the current value, or a delete if the row is absent. So events are trustworthy state.
  • With setResidencyById, a writer stores nothing for a row it doesn't own (its listener then gets delete). The sender never ships a live row to a peer that doesn't own it, but a deleted row's entry does cross in a base copy.
  • A base copy that carries any row of a table makes the receiver re-send the whole table to live subscribers (the reload re-snapshot, harper-pro#495).
  • Gap replay: each thread walks a database's audit log with one reusable iterator, skipped while the thread has no subscriber. A new subscription therefore receives every row written since the thread's last subscriber ended, delivered with the first commit after subscribing. It's harmless (idempotent), so subscribe-then-scan cannot miss a write. The size grows with the gap, which matters for a keeper restart.
  • An operations-API upsert stored rows on a non-owner in the harness (observed). The likely cause is that the component's residency function isn't installed on that thread. Production's own writes go through the component.

Next steps

  • Review and merge the harness (#216) and the two-node harness (#217).
  • Two-node test: the owner's subscription sees every replicated write; a join re-sends the whole table when the copy carries a row.
  • Queue-state response shape.
  • Keeper, claims from it, queue state, and the index dropped: #219, released as prerender-v0.93.0.
  • Churned, production-size run (1M rows with sustained reschedules): the load time decides how long a restart serves a partial order.
  • Console catch-up: #221, released as prerender-console-v0.18.0 (per-node keeper view and Health checks from queue-state, in flight from overview.leases, the floor and sweep panels removed).
  • Cluster queue state: a peer fan-out via util/peer.js, 503 on a missing peer or stale data, never a partial sum.
  • A batch-write arm in the harness that stays subscribed, so replay doesn't contaminate it.
  • Queue-state fixes found by the console catch-up (exact in-flight, unready before a node's first report, listsTruncated): #220, released as prerender-v0.93.1.
  • Split rows gained after a membership change from missed writes, in keeper_repaired, verify.repaired and exact. Today a node-list change reads as a subscription that missed writes.
  • A metric for lease refusals, including those from the bounded 8-slot probe window, which can refuse well below 100% occupancy. Today it is only a log warning.
  • Put carried on queue-state class rows. It is part of a class's identity, so without it two classes can look identical.
  • Expose a key's hold state (misses, held until) in explain, so a held key doesn't read as "overdue, not leased" with no reason.
  • A count of distinct held keys beside claim_wedged, which counts hold events.
  • keeper_live with a reason, and at a finer cadence than the backlog snapshot.

Open questions:

  • Do deleted entries persist in production's render_schedule? If they do, every base copy re-sends each receiver's whole table to the keeper. A deleted row crossing despite residency may be worth a harper-pro issue.
  • The keeper shares worker 0 with a replication socket (seen in the harness). A stalled worker 0 now degrades to serving its last set until it drains (each grant checked against its row); a stall longer than that still stops claims on the node.

Related: #80 (render priority), #218 (the store mutex).

🤖 Generated with Claude Code

Linguagem predominante
JavaScript
Estrelas
0
Forks
0
Merge médio
8h 48min
PRs com merge (30d)
58

Preparar o ambiente

Primeiros passos

  1. Leia a issue inteira e depois o guia de contribuição do projeto.
  2. Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
  3. Faça um fork do repositório e trabalhe em uma branch.
  4. Abra um pull request que referencie o número da issue.

Mais de HarperFast/prerender-plugin

Todas as issues de HarperFast/prerender-plugin

Issues semelhantes

Mais issues de JavaScript

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.