Queue state: per-node in-memory queue keeper, the render queue itself (built in #219)
Mantenedores costumam responder em até 1 dia
Ninguém assumiu esta issue ainda.
Avaliação
- Dificuldade
- 5/5
- Tempo estimado
- Mais de uma semana
- Facilidade para iniciantes
- 15/100
- Tipo de issue
- Funcionalidade
- Clareza
- Razoavelmente clara
- Status de atividade
- Estagnada
- Stack de tecnologia
- javascript
Direção de pesquisa
Start by reading PR #219, which contains the keeper, claims, queue-state response, and index removal described here. Review the linked harness work in #216 and #217 and the unchecked next steps; this issue is done only when the remaining production-scale, console, cluster-state, and harness work is addressed.
Escrita pelo modelo de indexação a partir do texto da issue.
Descrição
Summary
Expose the render queue's state cheaply, per node and cluster-wide, for the management console and as an autoscaling signal for the render fleet. The chosen design ("design C" below) keeps each node's queue in memory on worker 0, fed by a RenderSchedule subscription, and makes it the queue itself: RenderSchedule is its durable side and the nextRenderTime index is gone.
Status: merged in #219 and released as prerender-v0.93.0, with queue-state fixes in prerender-v0.93.1; not yet deployed. The claim floor, the ready-set sweep, the index claim scan and the index itself are removed in the same change, with no transitional fallback.
- Single node: a subscription adds about 46 µs of CPU per schedule write, the in-memory copy stays exact, and memory is about 200 B per row (about 50 MB per node at production scale).
- Two nodes: each node's keeper receives every replicated write and stays exact. Its cost is within round-to-round noise in process CPU, and adds 7–13 µs per write on the keeper's thread.
- Real Harper, upgrading 0.92.0 → 0.93.0 in place (single node, 2,001 rows): the index drop needed no manual step; the first claim 2 s after the restart was granted;
queue-statereported exactly the seeded due count; claims came in keeper order with no overlap; the verification walk repaired nothing. Rolling back rebuilt the index from the records, exact; on a multi-thread node it also needs a second restart once the rebuild finishes (harper-pro#681). Details in #219. - A base copy that carries any row of the table re-sends each receiving node's whole table to the keeper (about 1.2 s at 250k rows, fresh). The keeper absorbs it.
Not tested yet: a churned production-size store (it now decides only how long a restart serves a partial order), more than two nodes, a live membership change, and the peer check against harper-pro.
Queue state (GET /prerender_admin/queue-state, per node)
| Group | Fields |
|---|---|
now |
due (sitemap / discovered), in flight (leases), unclaimed (estimate), paused, status |
coming |
due in the next 15 / 60 minutes and 24 h; next 24 h by hour |
lateness |
due rows by lateness relative to cadence, split by sitemap vs discovered and by route; oldest due row per class. A single global head age would only measure how long the lowest-priority class has starved. |
flow |
per minute for the last hour: came due, added, triggered, rescheduled, removed |
trust |
live, phase, exact, state age, the keeper's load / publish / verification stats |
503 (with trust and the live now fields, no counts) whenever the keeper cannot vouch for its numbers. Not built yet: the cluster fan-out and the console view.
Options considered
| Change | Freshness | Notes | |
|---|---|---|---|
A. Publish what sweepReadySet already counted |
Small | Sweep interval (5 min) | Counts from the sweep's walk, capped at sweepCap. |
| B. Subscription-driven exact counters beside the old queue | Medium | Seconds | The first step of C. |
| C. In-memory queue per node (chosen) | Large | Seconds | The keeper decides claim order and refills the shared-memory ready set; the sweep, the claim floor and the index go. |
Design C, as built
- Load, serving from the first chunk. A node whose
system.hdb_nodesnames another node waits (up to 2 minutes) to see one, since ownership is not knowable before the node list is; a single node does not wait. Subscribe (listener form,omitCurrent: true), then walk the node's own rows by primary key in local chunks, publishing the ready set from what is held so far. Events apply as they arrive; a walked row is applied only if no event touched its key since the walk began. If an unreadable row stops the walk, the rest is read from the top down. - Hold
key → (due minute, class), and per classminute → Set(keys). A class is route × cadence (and whether it came off the row) × sitemap flag. Only rows this node owns by residency are held. - Stay current. Apply each
put/delete/invalidateevent by moving the key. With residency pinned by id, a node's subscription gets aputfor every write to a row it owns and a no-opdeletefor each write it makes to a row another node owns (measured, see Benchmark). - Serve claims. Every second, write the best few thousand keys into the shared ready set. A claim, on any worker, takes entries in order, point-reads each one's row locally, skips any no longer due, and leases the rest with an exclusive grant (no mutex; see #218). A key whose leases keep expiring with no result is held back in the lease table, 2 leases then 4, 8, … capped at its cadence. Until its first publish the node grants no claims and reports the queue status
unready; the keeper reports each change of status itself, so the fleet is woken by it. A stalled keeper is still served from until its last set drains. - Never reload once live. A route or default-interval change reclassifies held rows in memory; a membership change reclassifies (dropping rows no longer owned) and runs the verification walk (adding rows gained); a closed or failed subscription is reopened with backoff and walked.
- Check itself. Each publish re-reads the top 64 published rows (each key at most once a minute); a verification walk every hour compares every row with the table. Both repair what is wrong, and a repair never overwrites a newer event.
- Report state from the buckets, every 5 s, into a shared buffer any worker serves.
Priority order without rescoring. Rank is lateness relative to the row's own cadence, with a sitemap boost. Two rows in the same class share cadence and boost, so the one due earlier always ranks higher: each class list is already in priority order, and the best K is a merge over class heads.
What was given up with the index: ad-hoc due-time queries against the table. A rollback to 0.92.0 re-adds the index, which Harper rebuilds from the records (observed); 0.92.0's claims fail until the rebuild finishes.
Benchmark
Harness: bench/queue-keeper (#216); two-node harness in bench/queue-keeper/cluster (#217). The README has the method and caveats. Both run Harper 5.2.13, the version production runs, on a fresh corpus, so absolute costs are floors and ratios transfer.
Single node
harperfast/harper:5.2.13, 250k rows, 2 worker threads (writer + keeper).- 50k single-row writes per arm, 5 rounds, interleaved.
| No subscription | Subscription | Subscription + keeper | |
|---|---|---|---|
| put p50 / p99 | 92 / 183 µs | 97 / 188 µs | 98 / 212 µs |
| process CPU per write | 106 µs | 142 µs | 152 µs |
| keeper thread busy per write | — | 26 µs | 31 µs |
| events per write / lag p99 | — | 1 / 0.8 ms | 1 / 0.8 ms |
- Correctness: the keeper matched the table on 250,000 of 250,000 rows. A subscription on one thread sees every commit made on another.
- Memory: about 136 B/row of key strings plus 64–80 B/row of structures. That's 48 MB at 250k rows and 191 MB at 1M.
- Claim and state costs: top 5,000 due rows in 0.16–0.25 ms; full counts under 2 ms; rebuild 0.9 s for 250k rows (expect 20–30× on a churned store).
Two nodes
harperfast/harper-pro:5.2.13, two containers replicating over TLS, 2 worker threads each.- 250k rows per node, pinned the way production pins
RenderSchedule(setResidencyById, rendezvous hashing with production's hash). - Both nodes write rows of both owners at once: 50k single-row writes per node per arm, about 15k cluster writes/s. 3 rounds, interleaved.
| Per node | No subscription | Keeper |
|---|---|---|
| process CPU per cluster write | 134 µs (123–141) | 137 µs (133–138) |
| keeper thread (worker 0) busy per cluster write | 44–51 µs | 55–59 µs |
| put p50 / p99, own row | 131–135 / 267 µs | 126–132 / 240–247 µs |
- The owner's subscription sees every write replicated from a peer (a coverage check, 8 of 8 windows).
- Every distinct row written to a node, by either node, reached its keeper.
deletes matched the node's own writes to foreign rows exactly, and noputarrived for a foreign row.- The 1–4
puts per 50k that never arrived were superseded versions: Harper sends only a row's current version.
- Exact at the end: 500,000 of 500,000 rows, each on its owner, and each keeper equal to its own table.
- Lag, writer's commit to the owner's listener: replicated p50 0.3–0.4 ms, p99 42–105 ms at 15k cluster writes/s.
- Cost:
- Process CPU, paired per round (keeper minus no subscription): −4 to +12 µs per cluster write, a spread as large as the effect.
- The keeper's thread rose consistently, by 7–13 µs per cluster write, which is 10–17 µs per delivered event against 31 µs on a single node.
- The keeper's own work was 4.5–4.8 µs per
put.
- The keeper shares its thread with replication: the table's replication socket was on worker 0 on both nodes.
What a join costs (the reload burst). add_node requested a full copy each way, and row counts never changed. A live keeper's experience depended only on whether a copy carried any row of the table (12 runs):
- The trigger was one deleted row. Its entry crossed despite residency, and the receiver skipped storing it. Having received a row, though, the receiver wrote a
reloadmarker, and Harper re-sent every live subscriber the node's whole table. At 250k rows that took about 1.2 s per node, and 0.4 s of keeper work. - With no row in the copy, nothing was re-sent.
- A commit on the keeper's thread made no difference.
- In production (inference):
RenderSchedulerows get deleted, so expect every base copy ofrender_scheduleto re-send each receiver's table. Base copies happen on a join, a clone, an interrupted first copy, or a peer falling behind the retained audit log.
Harper behaviour this relies on (5.2.13 source, confirmed by the runs)
- The subscription work runs on the subscribing thread; commits are coalesced via
setImmediate. - Harper re-reads the row before delivery, drops stale versions, and sends the current value, or a
deleteif the row is absent. So events are trustworthy state. - With
setResidencyById, a writer stores nothing for a row it doesn't own (its listener then getsdelete). The sender never ships a live row to a peer that doesn't own it, but a deleted row's entry does cross in a base copy. - A base copy that carries any row of a table makes the receiver re-send the whole table to live subscribers (the
reloadre-snapshot, harper-pro#495). - Gap replay: each thread walks a database's audit log with one reusable iterator, skipped while the thread has no subscriber. A new subscription therefore receives every row written since the thread's last subscriber ended, delivered with the first commit after subscribing. It's harmless (idempotent), so subscribe-then-scan cannot miss a write. The size grows with the gap, which matters for a keeper restart.
- An operations-API
upsertstored rows on a non-owner in the harness (observed). The likely cause is that the component's residency function isn't installed on that thread. Production's own writes go through the component.
Next steps
- Review and merge the harness (#216) and the two-node harness (#217).
- Two-node test: the owner's subscription sees every replicated write; a join re-sends the whole table when the copy carries a row.
- Queue-state response shape.
- Keeper, claims from it, queue state, and the index dropped: #219, released as prerender-v0.93.0.
- Churned, production-size run (1M rows with sustained reschedules): the load time decides how long a restart serves a partial order.
- Console catch-up: #221, released as prerender-console-v0.18.0 (per-node keeper view and Health checks from
queue-state, in flight fromoverview.leases, the floor and sweep panels removed). - Cluster queue state: a peer fan-out via
util/peer.js, 503 on a missing peer or stale data, never a partial sum. - A batch-write arm in the harness that stays subscribed, so replay doesn't contaminate it.
- Queue-state fixes found by the console catch-up (exact in-flight,
unreadybefore a node's first report,listsTruncated): #220, released as prerender-v0.93.1. - Split rows gained after a membership change from missed writes, in
keeper_repaired,verify.repairedandexact. Today a node-list change reads as a subscription that missed writes. - A metric for lease refusals, including those from the bounded 8-slot probe window, which can refuse well below 100% occupancy. Today it is only a log warning.
- Put
carriedonqueue-stateclass rows. It is part of a class's identity, so without it two classes can look identical. - Expose a key's hold state (misses, held until) in
explain, so a held key doesn't read as "overdue, not leased" with no reason. - A count of distinct held keys beside
claim_wedged, which counts hold events. -
keeper_livewith a reason, and at a finer cadence than the backlog snapshot.
Open questions:
- Do deleted entries persist in production's
render_schedule? If they do, every base copy re-sends each receiver's whole table to the keeper. A deleted row crossing despite residency may be worth a harper-pro issue. - The keeper shares worker 0 with a replication socket (seen in the harness). A stalled worker 0 now degrades to serving its last set until it drains (each grant checked against its row); a stall longer than that still stops claims on the node.
Related: #80 (render priority), #218 (the store mutex).
🤖 Generated with Claude Code
- Linguagem predominante
- JavaScript
- Estrelas
- 0
- Forks
- 0
- Merge médio
- 8h 48min
- PRs com merge (30d)
- 58
Preparar o ambiente
- Sem Dockerfile nem arquivo Docker Compose
- Sem modelo de pull request
- Ler o guia de contribuição
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de HarperFast/prerender-plugin
-
bug
Dificuldade 3/5 1-2 dias Facilidade para iniciantes 74/100
HarperFast/prerender-plugin#218 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 4/5 3-5 dias Facilidade para iniciantes 48/100
HarperFast/prerender-plugin#189 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 5/5 Mais de uma semana Facilidade para iniciantes 35/100
HarperFast/prerender-plugin#185 · 1 comentário ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 5/5 Mais de uma semana Facilidade para iniciantes 25/100
HarperFast/prerender-plugin#183 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 5/5 Mais de uma semana Facilidade para iniciantes 35/100
HarperFast/prerender-plugin#180 ·
Mantenedores costumam responder em até 1 dia
Todas as issues de HarperFast/prerender-plugin
Issues semelhantes
-
Add google analyticsAberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 68/100
NCAR/music-box-interactive#628 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 68/100
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 65/100
remotion-dev/remotion#11847 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 86/100
phoenixframework/phoenix_live_view#4456 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 88/100
AllTheMods/ATM-10#4436 ·
Mantenedores costumam responder em até 5 dias