Raw-document caching per route (shipped in #181); probe-gating superseded; size post-processing and deploy detection still open
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 35/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- javascript
- Domain
- backend, performance
Research direction
Start with the existing postProcess.removeAttributes path and the @harperfast/prerender-browser build for non-DOM raw-document processing. For deploy detection, read actOnTrip, invalidateScope, requestSweepReseed, and the canary flow; done should cover bounded size post-processing and correlated bulk invalidation while preserving the stated no-op defaults and avoiding template-trip reseeding.
Written by the indexing model from the issue text.
Description
Status
The raw-document half of this is implemented in #181. The probe-gating half turned out not to need building. What remains is two narrower things, at the bottom.
Superseded: render.gateOnProbe is a no-op you can already configure
The original proposal was a gate: skip a scheduled render when the probe's signature is unchanged, with maxInterval as a backstop.
That produces exactly the same renders as setting the route's renderInterval to the backstop and letting the existing probe trigger pull renders forward. Same change-driven renders, same backstop, same keys × r / (1 - (1-r)^b) arithmetic — and zero new code on the claim path.
It is also already the running behaviour on the deployment this was written for: the probe is armed, anchored to the origin's nightly change, and triggerRevalidate hard-expires and re-files a page the day its watched fields change.
Two reasons not to build the gate as specified, beyond redundancy:
- It inverts the safety property the probe is documented to have.
configSchema.jsstates it plainly: the probe is "an accelerator on top of the baseline render cadence, never a gate on it — the failure mode to survive is the origin replatforming under a rule." A gate turns a stale rule from noisy over-rendering into a silent freeze. - It would cost a per-row read on the hot path. The ready-set sweep already keeps a per-sweep memo to avoid so much as a URL parse; a
ProbeStatepoint read per due row inside that cursor loop is the version to refuse.
What would have been worth building is the part the proposal did not have: a coverage guarantee. Raising an interval to a long backstop silently makes a probe outage, a dropped rule match, or an un-seeded URL into a multi-day staleness event. If this is ever revisited, the shape is a grant evaluated at schedule-write time — the long interval applies only where positive probe coverage exists (enabled, not dry-run, a rule matches, the pass is not cycle-behind) — resolved inside resolveEffectiveInterval, which every schedule writer already passes through. Never a per-row check in the sweep.
Shipped in #181: raw-document caching
Per-route rawCache: true plus render.raw.*. On a miss for an opted-in route the body already being streamed to the crawler is teed, captured to a byte cap, and stored — so the next crawler asking for that URL is a cache hit instead of another origin round trip. No render capacity, no second origin request, one write.
Differences from the sketch in the original proposal, all deliberate:
- Its own table in its own database, not
PrerenderedPage. That table has no expiration and this population is unbounded and long-tailed by construction; it is not residency-pinned, so every row would replicate to every node; andpage_cacheis the database whose transaction log has been measured pegging worker threads at 85–95% of CPU on the receive side of a copy. - It is a serve-path tee, not a scheduled fetcher. A
policy: rawthat schedules raw refreshes would add a speculative origin fetch per URL per TTL for pages nothing crawls — worse than the misses it replaces. Teeing is demand-driven by construction and converts exactly the misses that already exist. - It only ever replaces an origin proxy. A stored document is read when
PrerenderedPageheld nothing at all. A stale or invalidated row means a render is coming, and the live origin beats a stored copy of it. skipwas not implemented because it already exists:discoverTargets: falsegates a route out of the render rotation, and it is the intended pairing forrawCache.
Still open
1. Post-processing raw documents for size
Raw origin documents can be much larger than a rendered snapshot — frameworks that hydrate from serialised props put the majority of the bytes in attributes no consumer reads, and some documents exceed search engines' soft size limits.
#181 caps and refuses oversize documents rather than shrinking them, and the cap is on compressed bytes (the body is stored in the origin's own encoding so the serve path can re-encode from the stored header). So the cap bounds memory and storage; it is not a test against a decompressed size limit.
The existing postProcess.removeAttributes cannot be reused: it runs as document.querySelectorAll in page context, so it needs a browser. A non-DOM streaming rewrite belongs in @harperfast/prerender-browser, which has a build step and where a parser dependency is cheap — not in the plugin, which runs inside Harper.
2. Detecting a front-end deploy
A deploy changes every page's content-hashed assets, and a snapshot that outlives the origin's retention of its assets renders unstyled to bots. The original proposal asked for a documented manual trigger for bulk invalidation.
A better shape is available and cheaper: a deploy is a correlated change, so it does not need a per-URL walk to find — it needs one detection and one bulk invalidation. That is exactly the canary's existing shape (invalidateScope on a trip). A second canary cohort that fetches documents and digests a template-only projection — <script src>, <link href>, asset hashes, explicitly excluding content — would catch a deploy within its interval for a few hundred fetches a day, reusing machinery that already exists and already knows how to invalidate a scope.
It is also unusually well conditioned as a detector: quiescent is ~0% of the cohort and a deploy is ~100% simultaneously, so there is a wide safe band for the threshold — unlike a content canary, where the natural change rate and the event rate can sit close enough together to produce false trips.
Two couplings to get right if this is built:
- A template trip must not reseed.
actOnTripfiresrequestSweepReseedunconditionally, which is right for a mass content change but would burn a full corpus probe pass rewriting content baselines a deploy did not touch. - A trip should file schedule rows, not just an epoch. Invalidation demotes but never reschedules, so healing depends on cadence — and the longer the route's interval, the longer the whole scope serves from origin after a trip.
Constraints (unchanged)
- Every default must be an exact no-op. Unset config produces byte-identical behaviour, the way
config.overridesdoes. - Nothing here may assume a particular site's routes, change rate, or SSR completeness.
Related
- #181 — raw-document caching (implements the raw half of this)
- #182 —
explainreports the full cadence resolution, which is how "what interval is this URL actually on" stopped being guesswork - #179 — per-device schedule rows folding
🤖 Generated with Claude Code
- Dominant language
- JavaScript
- Stars
- 0
- Forks
- 0
- Avg merge
- 9h 10m
- Merged PRs (30d)
- 56
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from HarperFast/prerender-plugin
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
HarperFast/prerender-plugin#176 ·
-
Difficulty 4/5 3-5 days Newbie friendliness 48/100
HarperFast/prerender-plugin#189 ·
-
Difficulty 5/5 Over a week Newbie friendliness 35/100
HarperFast/prerender-plugin#185 · 1 comment ·
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
HarperFast/prerender-plugin#183 ·
-
Difficulty 4/5 3-5 days Newbie friendliness 38/100
HarperFast/prerender-plugin#179 ·
All issues in HarperFast/prerender-plugin
Similar issues
-
documentation
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
githubnext/gh-aw-workshop#3692 ·
-
agent/guide documentation hive/hosted-available-lke648397-260827-5n31
Difficulty 2/5 1-3 hours Newbie friendliness 90/100
-
Add: BuyPass TV Openchannels:add check:passed
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
-
S: triage
Difficulty 1/5 Under an hour Newbie friendliness 85/100
-
bug
Difficulty 1/5 Under an hour Newbie friendliness 90/100
apache/cloudstack#14222 ·