Capture memory is bounded in uncompressed bytes: bound the crawler's queue and compress as the capture reads, so the cap means stored size
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 38/100
- Tipo de issue
- Nueva funcionalidad
- Claridad
- Bastante claro
- Estado de actividad
- Activo
- Stack tecnológico
- javascript
- Área
- backend, performance
Línea de trabajo
Start in util/rawCache.js, especially teeForCapture and collectBody, and review the stalled-reader test described in #229. First resolve the queue-wait versus capture-abandon and option-naming questions, then verify the stated done conditions: bounded memory, a 1.5 MB uncompressed 404 stored under a 512 KB cap, and a failing revert check without the queue bound.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
The raw and negative caches' maxBytes counts a body as received. So an origin that sends error pages uncompressed needs a cap sized for the uncompressed body, though the copy that is stored is gzipped 7-13x smaller (#229). This issue is about making the cap bound what is held and stored, not what crossed the wire.
Problem
- Why the cap counts received bytes. It is what bounds a capture's memory: about 2x
maxByteseach.- One copy is the captured chunks.
- The other is what
tee()keeps for the crawler's unread branch, because the capture reads in a tight loop and runs ahead (teeForCapture/collectBodyinutil/rawCache.js; measured in #229).
- What it costs, on one production origin (observed, 24h of
prerender_ops negative_cacheover 4 nodes):- Its retired-product 404s arrive uncompressed at 190 KB-1.5 MB.
- At a 1 MB cap, 6.5k were refused as
oversize, about 3.4% of the 404s stored. These are recorded counts; Harper's analytics record about 2/3 of real. - Each would have stored at roughly 80-210 KB (inferred from the 7-13x gzip ratio).
- The workaround. The deployment raised
render.negative.maxBytesto 2 MB. That raises the per-capture memory bound to about 4 MB, to admit bodies that store at about 150 KB. It is affordable there only because few captures run at once.
Why compressing the capture alone is not enough
#229 considered capping on compressed bytes and kept the received-bytes cap for this reason:
- Gzipping the captured branch as it reads shrinks only the captured half.
- The tee would still hold everything the crawler has not read yet, uncompressed.
- With the cap on compressed bytes, a slow crawler could make one capture hold about 10x the cap in that other half.
Proposal
- Bound the crawler's queue separately from
maxBytes. Put a high-water mark (say 256 KB) on the unread bytes held for the crawler's branch.ReadableStream.tee()has no such limit, so this needs a small custom tee. Past the mark, one of two things happens:- a. The capture waits for the crawler. This restores backpressure to the origin, but the capture slot is held for as long as the crawler takes to read.
- b. The capture is abandoned, counted under a new outcome such as
capture-slow.
- Compress the captured branch as it reads when the origin sent no encoding (zlib's streaming gzip, level 6, which runs on the thread pool). Count the cap on the compressed output. A body the origin already encoded is captured as it is today.
- Result: per-capture memory is about high-water + compressed bytes + deflate state (~256 KB at zlib defaults). It no longer depends on the uncompressed size, so the cap can be a stored-size cap.
Open questions
- 1a or 1b? Slot occupancy suggests waiting is affordable: in the same 24h, raw
capture-busy4 against ~858k stored, negative 0. This is a hypothesis; the share of slow crawlers is unmeasured. - Keep
maxBytesor add an option? RedefiningmaxBytesto mean compressed bytes would silently let an existing config admit bodies ~10x larger. A new option, for examplemaxStoredBytes, alongside a queue high-water option, is safer. - Can the store step be dropped?
storableBodygzips once on the detached store path today. With compression at capture time it would only pass bodies through, but rows stored before the change still need the re-encode on the serve path.
Done when
- A stalled-reader test, like the one in #229, shows memory per capture bounded independently of the body's uncompressed size.
- A 1.5 MB uncompressed 404 is stored under a 512 KB cap. A deployment that raised the cap to admit such bodies can lower it again.
- A revert check: with the queue bound removed, the stalled-reader test fails.
🤖 Generated with Claude Code
- Lenguaje dominante
- JavaScript
- Estrellas
- 0
- Forks
- 0
- Merge medio
- 8 h 11 min
- PR fusionados (30 d)
- 73
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de HarperFast/prerender-plugin
-
Dificultad 5/5 Más de una semana Aptitud para principiantes 25/100
HarperFast/prerender-plugin#245 ·
Los mantenedores suelen responder en 1 día
-
enhancement
Dificultad 5/5 Más de una semana Aptitud para principiantes 25/100
HarperFast/prerender-plugin#244 ·
Los mantenedores suelen responder en 1 día
-
enhancement
Dificultad 5/5 Más de una semana Aptitud para principiantes 30/100
HarperFast/prerender-plugin#242 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
enhancement performance
Dificultad 5/5 Más de una semana Aptitud para principiantes 38/100
HarperFast/prerender-plugin#233 ·
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 3/5 1-2 días Aptitud para principiantes 74/100
HarperFast/prerender-plugin#218 ·
Los mantenedores suelen responder en 1 día
Todos los issues de HarperFast/prerender-plugin
Issues similares
-
Edit: RTE News LogoAbiertocheck:failed logos:edit
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
iptv-org/database#36354 · 1 comentario ·
Los mantenedores suelen responder en 4 días
-
agentic-workflows documentation workflow-editor
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
githubnext/gh-aw-workshop#4139 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
CircuitVerse/CircuitVerse#7967 · 1 comentario · 1 reacción ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
CopilotKit/aimock#491 ·
Los mantenedores suelen responder en 1 día
-
cvss-severity:high devguard l3montree-cybersecurity/.../devguard-documentation pkg:devguard/l3montree-c.../devguard-documentation risk:low state:open
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
l3montree-dev/devguard-documentation#338 · 1 comentario ·