Rule proposal: long-context-redact — collapse over-budget threads on long pages
I maintainer di solito rispondono entro 3 giorni
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 35/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Tranquilla
- Stack tecnologico
- typescript
- Ambito
- security
Direzione di ricerca
Start with rules.md and compare the existing comments-redact, reviews-redact, cross-origin-frame-redact, and irrelevant-sections-redact rules referenced in the proposal. Determine the rule entry point, configuration and placeholder behavior from those implementations, then define tests for per-section limits, reveal-on-demand behavior, denylist exceptions, and structural significance markers.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Category
New defense rule
What problem does this solve?
Even when a page is "clean" (no injection, no dark patterns), sheer length is a defense surface. Liu et al. (TACL 2023), Lost in the Middle: How Language Models Use Long Contexts, document a U-shaped accuracy curve: LLM retrieval and reasoning degrade sharply when relevant information sits in the middle of a long context window, even on models advertised as long-context.
A page with 200 comments or 80 reviews pushes the agent's task instructions and the page payload to the edges, where the agent both (a) misses the answer and (b) is more susceptible to mid-context injection that benefits from positional dilution. The existing comments-redact and reviews-redact rules remove these surfaces entirely. But when the comments are the task ("summarize the top criticism of this product"), removal is wrong — the agent needs the content, just not all 800 entries.
Proposed solution
On pages whose visible text exceeds a budget (proposed: 50k chars; tunable), collapse:
- Comment threads past the first N entries
- Review lists past the first N entries
- Reply chains past the first N levels of depth
into the same click-to-reveal placeholder shape cross-origin-frame-redact, comments-redact, and irrelevant-sections-redact already use. Agents that need the tail can reveal explicitly; default behavior preserves the head of the list (typically highest-quality on engagement-sorted platforms).
Alternatives considered
- Lower the cap in
comments-redact/reviews-redact. Doesn't help — those rules are all-or-nothing today; this proposal is the "keep the head" variant. - Reader-mode extraction (Readability.js). Already part of the prior-art lineage; Readability picks a main article and discards comments wholesale. We want the keep-some-comments shape.
- Generic LLM trim. Same shape as
irrelevant-sections-redact, but the trigger here is length, not engagement-rail recognition — no LLM call needed for a "keep first N" heuristic.
Controlling false positives
- Default-off. Until per-host hit/skip data confirms the head-N preserves task-relevant content, this rule should ship default-off — same posture as
irrelevant-sections-redact. Users opt in when their workflow tolerates the trade-off. - Per-host denylist. Sites where the tail is structurally load-bearing — Hacker News (highly-rated child comments outvalue top-level), GitHub issue threads (resolution often in the last reply), Reddit (
AskScience-style threads with cited replies), public-comment portals like regulations.gov, court records — never apply the rule. - Preserve elements with structural significance markers. Reddit "best answer" flags, Stack Overflow "Accepted" badges, GitHub "Marked as answer",
aria-label*="solved", and rows with engagement scores above a per-host percentile. The head-N count should not blindly drop a flagged answer because it sits at position 47. - High length threshold. 50k visible chars is roughly 12k tokens — well above the budget where Lost-in-the-Middle starts to bite for current frontier models. Tuning low risks redacting pages that fit comfortably in context.
- Reveal-on-demand contract. Collapsed regions become click-to-reveal placeholders, not deletions. An agent that detects the placeholder shape can decide to expand — same affordance as
cross-origin-frame-redact. - Per-section quotas, not page-wide. Apply N independently to each thread/list so a page with multiple distinct conversation surfaces (an article with comments + a sidebar of reviews) doesn't lose representation in one to keep room for the other.
- Skip when the page is the agent's task surface. If the user's task implies "summarize all reviews", the agent can't tell us — but a per-host denylist for the platforms where this is the common ask (review-aggregator sites like Trustpilot, Amazon SERP review pages once the user navigates to "see all reviews") gets most of the way there.
Prior art / references
- Liu et al. (TACL 2023). Lost in the Middle. https://arxiv.org/abs/2307.03172 — foundational motivation. Code/data: https://github.com/nelson-liu/lost-in-the-middle.
- Mozilla Readability.js (Apache 2.0) — boilerplate prior art already cited in
rules.md. - Kohlschütter et al. (WSDM 2010) — same lineage.
Tagged Impact M / Complexity H.
- Lingua principale
- TypeScript
- Stelle
- 34
- Fork
- 3
- Merge medio
- 3g 3h
- PR unite (30g)
- 35
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di pixiebrix/agent-browser-shield
-
enhancement question
Difficoltà 5/5 Più di una settimana Idoneità per principianti 30/100
pixiebrix/agent-browser-shield#178 ·
I maintainer di solito rispondono entro 3 giorni
-
Rule proposal: canvas-text-annotate — flag canvas/video text surfaces invisible to DOM walkersApertaenhancement rule-proposal
Difficoltà 5/5 Più di una settimana Idoneità per principianti 28/100
pixiebrix/agent-browser-shield#123 · 1 commento ·
I maintainer di solito rispondono entro 3 giorni
-
enhancement rule-proposal
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
pixiebrix/agent-browser-shield#122 ·
I maintainer di solito rispondono entro 3 giorni
-
enhancement rule-proposal
Difficoltà 5/5 Più di una settimana Idoneità per principianti 38/100
pixiebrix/agent-browser-shield#121 · 1 commento ·
I maintainer di solito rispondono entro 3 giorni
Tutte le issue di pixiebrix/agent-browser-shield
Issue simili
-
level/task reporter/qa type/bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
wazuh/wazuh-dashboard-plugins#9310 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
cybersemics/treecrdt#267 ·
-
Difficoltà 1/5 1-3 ore Idoneità per principianti 85/100
wiz-sec-public/backstage-plugin-wiz#16 · 1 commento ·
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
solana-foundation/solana-com#2245 ·
I maintainer di solito rispondono entro 1 giorno