Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Operator outcomes: what prerendering and raw caching are for, how a site sees they work, and what to change

Aperta
#244 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
5/5
Tempo stimato
Più di una settimana
Idoneità per principianti
25/100
Tipo di issue
Funzionalità
Chiarezza
Da chiarire
Stato di attività
Attiva
Stack tecnologico
javascript
Ambito
observability

Direzione di ricerca

Start with README.md, METRICS.md, and the existing traffic, queue, health, probe, corpus, and invalidations console views. Read the current KPI and metric paths, including util/demandLadder.js, util/documentFacts.js, and resources/analytics/write.ts, before resolving the fidelity, waste, and raw-cache questions. Done means an agreed scorecard design and implementation plan, with route-level verdicts and recommendations, approved before code is built.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

enhancement

Status: proposed 2026-10-02, not started. This is a handoff: the agent who picks it up owns the research, the design and the build, with the maintainer's agreement before code. The current entity and freshness work (#239, #240, #241, #243) is being deployed by another agent. Don't touch those branches; build on main after they merge.

TL;DR

  • What it is for. A site runs prerendering and second-tier raw caching so that every crawler, answer engine and preview bot gets a complete, current, fast copy of each page, without that traffic landing on the origin.
  • What is missing is the operator's view of that. Today there are 253 config options and dozens of metric series. No screen answers the two questions a customer actually has: is it doing its job, and what should I change.
  • The ask: work out, from the customer's side, what they need to see. Then build it as a small scorecard: one row per outcome, each with one headline number, a verdict (good / watch / bad) and the lever it points to. Every row sliced by route.
  • Two gaps to close first:
    • a fidelity number: is the snapshot complete, and equivalent to what users get;
    • a waste number: renders spent on pages no crawler asked for.
  • Plus one recommendation per route: render, raw-cache or proxy.
  • Research the real outcomes our numbers stand in for: indexing, crawl rate, merchant-feed approvals, answer-engine accuracy.

1. Who uses it, and what for

Who What they need from it What hurts them
SEO / growth lead Crawlers and answer engines see every page complete and current. Indexing, rich results and merchant listings are not harmed. A snapshot missing content. A stale price or stock state (a merchant feed checks the landing page against the feed). A wrong status or canonical. Anything that reads as cloaking.
Platform / SRE Bot traffic and its surges kept off the origin. No incidents caused by the cache or by our own renders and probes. Bot surges reaching the origin. Our renders and probes adding origin load. Serving errors. A personalized document stored and replayed to everyone.
Capacity / budget owner Render spend matched to what bots use. Knowing when to scale up or down. Rendering pages bots never ask for. Too little capacity, which shows as staleness. Too much, which shows as waste.
Operator tuning it Which knob to turn, and proof afterwards that it worked. 253 options without a map from question to knob.

"Bots" is broader than search.

  • Search engines: they may render JavaScript themselves, under a budget and with a delay.
  • Answer and generative engines: many do not execute JavaScript, so this is how they see the content at all. Hypothesis, per engine; worth measuring.
  • Link-preview bots.
  • Tools acting for a person. On one production deployment, a read-aloud agent fetching pages for users was 0.72% of bot serves over 24h, served from both caches (observed, 2026-09-27). So a stale or broken copy can reach people too.

Second-tier raw caching is the other half. It stores the origin's own HTML, for:

  • routes whose server-rendered document is already complete for crawlers;
  • the long tail beyond render capacity;
  • 404/410 answers (the negative cache).

It gives offload with no render cost. Its risks are personalization (has-cookie, assumeShared), staleness, and content that only exists after JavaScript runs. So the customer's per-route decision is: render, raw-cache, or proxy.

2. The outcomes, and the one number each

# Outcome Headline number (per route) Today
1 Answered from cache, fast: offload, response time, and through those crawl rate Cache-served share; coverage misses (misses the origin could have served); hit vs origin serve time Have it: the Traffic view's KPI row, bot_serve, route_serve.
2 Rendered where rendering matters Share of serves that are a rendered snapshot, a raw document, or the origin Have the mix. Gap: nothing says, per route, whether raw is good enough.
3 Fresh: as up to date as possible Fresh hits (inside the cadence); staleness against the cadence; served wrong, meaning how many served copies were found to differ from the origin, for how long, and in which field Have it: the KPI row, plus served_wrong (0.102.0, #243).
4 Faithful: complete, and equivalent to what users get Share of renders whose readiness contract held; snapshots missing their hydration marker Gap. Readiness verdicts are emitted in report-only mode, but their panel is #189, not built. Parity tooling exists only offline.
5 Correct for indexing Served status matches the origin's; no page served whose canonical names another URL; noindex respected Partial: the status-code and render-outcome panels, and the entity-serve guards. There is no single number.
6 Worth it, and right-sized Renders per crawler serve; renders of pages no crawler requested within N days; wasted renders (verdicts, duplicate spellings); CPU-seconds per render; net offload Partial: net offload, render time and outcomes. Gap: "renders nobody asked for".
— Guardrails Origin load we add (net offload already subtracts our own renders, probes and sitemap fetches); 5xx and blob faults served; the cost of observability itself Have it, scattered across Health and Traffic.

Why fidelity is listed explicitly: before browser v1.15.0, every deployed page was cached without being hydrated, while the render reported 200, non-empty, indexable, with zero failures (verified, this repo's CLAUDE.md, "Hard-won lessons"). Every number in rows 1–3 looked healthy. Fast, fresh and wrong is the failure the current screens cannot see.

Why waste is listed explicitly: on one deployment, canonical-mismatch was 8.2% of all render outcomes (verified, the util/entityGate.js header). Those are renders that produced nothing to serve.

3. The decisions a customer makes, and what should drive each

Decision Signal Lever
More render capacity? Queue lateness and change-to-cache lag rising together with staleness, while the waste number is low Render nodes. Fix waste first.
Add pages? Coverage misses by cause, with how often each URL is asked for again (crawl breadth) A route, a sitemap, discovery. For duplicate spellings, the entity serve.
Remove pages, or slow them down? Renders nobody asked for (gap). The demand ladder does not slow cold pages: a route's base cadence is its ceiling, so a page no crawler visits renders at base indefinitely (util/demandLadder.js, effectiveLadder; the render.demand config text says otherwise and is wrong) Drop routes or sitemaps; lengthen the route's base renderInterval.
Change cadence or TTL? served_wrong per field, compared with staleness. Stale but rarely wrong: render less often. Often wrong on price or stock: shorter cadence, or the probe and serve-time checks. Route renderInterval, swrTtl, changeProbe.*.
Render, raw-cache or proxy a route? Whether the origin's document is equivalent to the rendered page for crawlers (gap), against the render cost rawCache per route, the route's mode.
Arm a feature? Its dry-run number: would-serve, would-adopt, would-gate, the negative cache's would-serve-live The feature's switch.
Something changed: is it us or the site? A site release changing templates or copy (gap: on one deployment one went unnoticed for days, observed); a fidelity drop; a served-wrong spike Re-render, readiness contracts, an incident.

4. Design principles

  • Few numbers, each with a verdict and its lever. Following the console's existing preference: dense, numbers first, prose behind the "?" toggle. An overview earns its place only as "what's worrying right now", never as a smaller copy of other views. A "watch/bad" verdict is a Health check with an explicit threshold.
  • Per route. ingress.routes is the partition axis. Named page types were tried and dropped on purpose; don't revive them.
  • Every switch has a dry-run number first. That pattern already exists; keep it uniform.
  • Measure at detection time, not per request. From Harper's analytics source (resources/analytics/write.ts, 5.0.28, verified):
    • a boolean value costs a counter increment;
    • a numeric value costs a Float32Array slot, plus a sort of that combination's values at every one-second flush, per worker.
    • So do not add a per-request numeric series without a strong reason; served_wrong is emitted once per detection.
  • Metric pruning is someone else's. The maintainer has asked for it as separate work. Don't remove series here, beyond what a scorecard row replaces one-for-one.
  • Proxies against the truth. Our numbers stand in for indexing, crawl rate, merchant-feed approvals and answer-engine accuracy. Say which one each row stands for.

5. Open questions to research first

  1. Fidelity, cheaply, in production. Candidates:

    • the share of renders whose readiness contract held (exists, report-only);
    • a framework's hydration marker (site-specific, so host-supplied code under #242, not config);
    • a sampled comparison against a user-agent render;
    • render outcome mix.

    Which one would have caught the v1.15.0 failure?

  2. Waste. Can the demand tracker (Bloom slices of which URLs bots visited) be joined with render completions to give "rendered, not requested in N days", without a per-URL write?

  3. Raw or render, per route.

    • util/documentFacts.js can already read a raw document's facts (title, canonical, offers, breadcrumbs) and compare them with a rendered page's.
    • Is a sampled agreement rate a sound recommendation? At what threshold?
    • What does it miss that only shows after JavaScript runs?
  4. External outcomes.

    • Search Console reports crawl rate, response time and indexing.
    • Merchant Center reports disapprovals.
    • Answer engines' citations can be sampled.

    Should the console show any of these? They need the customer's credentials. Or should it document how to correlate?

  5. A site change. The template-census idea: hash the page's asset names to get a release id, census boilerplate text and structure per route, over origin documents and snapshots. Comparing the two tells whose change it was. Is this the right mechanism, and how cheap can it be?

  6. Per crawler. Which bots matter differs by customer (search, answer engines, previews). Should the scorecard be per crawler class as well as per route?

  7. Verdict thresholds. What is "good" for each row on a typical site, and should it be configured or derived from the site's own history?

6. Constraints

  • This repo is public. No customer names, hostnames or IPs: say "one production deployment".
  • Fix forward, with no transitional fallbacks. Each phase replaces what it supersedes.
  • The plugin, console and render fleet ship separately. A new series needs the console to chart it, or waive it with a reason. The coverage test in packages/console/test/adminAssets.test.js enforces that.
  • The bot path runs at crawler volume. Nothing on it may allocate, read or emit per request without a measured reason.

7. How to run it

  1. Survey. Read README.md, METRICS.md, the console views (traffic, queue, health, probe, corpus, invalidations), and the issues linked here.
  2. Ask the maintainer which customer decisions matter most and which external outcomes they can reach.
  3. Propose. Post the scorecard as a comment on this issue: rows, the number behind each, verdict thresholds, the lever, and which current panel each row replaces or links to. Rewrite this body once it's agreed.
  4. Build in phases, each shippable:
    1. the scorecard from series that already exist;
    2. the fidelity number, with #189;
    3. the waste number;
    4. the raw-or-render recommendation per route;
    5. external outcomes, if agreed.

Related

  • #242: the plugin redesign. The scorecard is the operator-facing half of what "exemplary" means there.
  • #189: the readiness panel, and the start of the fidelity number.
  • #243: served_wrong, the start of the freshness row's "how wrong, for how long".
  • #237 and #166: entity serve and the entity registry, which feed the coverage and correctness rows.

🤖 Generated with Claude Code

Lingua principale
JavaScript
Stelle
0
Fork
0
Merge medio
8h 11m
PR unite (30g)
73

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di HarperFast/prerender-plugin

Tutte le issue di HarperFast/prerender-plugin

Issue simili

Altre issue su JavaScript

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.