[Bug]: Post-navigation page operations have no timeout, and page cleanup is skipped on task cancellation — pages can leak and crawls can hang forever
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 52/100
Línea de trabajo
Start in async_crawler_strategy.py at the post-navigation operations and the cleanup finally block, then inspect browser_adapter.py and browser_manager.py for timeout behavior. Reproduce the busy-page hang and cancellation path described in the issue. Done means post-navigation work is bounded by page_timeout and cancelled tasks still close their pages.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Summary
Found while investigating #2202 (Docker containers accumulating renderer processes for weeks). Two related defects in async_crawler_strategy.py make it possible for a crawl to hang indefinitely and for its page to never be closed.
1. Nothing after navigation has a timeout
page_timeout only bounds page.goto() (async_crawler_strategy.py:762-764). Every page interaction after navigation is an un-timed call into the page's JS engine:
- overlay/consent removal —
remove_overlay_elements(:1550) - body-visibility check,
css_selectorextraction (:1102), image-dimension updates page.content()(:1115)- shadow-DOM flattening, iframe processing, user
js_code
The adapter is a plain pass-through to page.evaluate (browser_adapter.py:61-65), which has no timeout in Playwright. The repo never calls set_default_timeout except inside if self.config.accept_downloads: (browser_manager.py:1215-1217) — and default timeouts would not cover evaluate anyway. Only scan_full_page is wrapped in asyncio.wait_for.
A page whose main thread goes busy after DOMContentLoaded (bad loop, broken ad script, hostile page) therefore hangs the crawl forever, holding its page (= one renderer process) open. Verified with a page that starts a busy loop right after DCL: navigation succeeds, then arun() never returns, far past page_timeout.
2. The page-closing finally doesn't survive cancellation
The cleanup block (async_crawler_strategy.py:1216-1231) guards with except Exception. asyncio.CancelledError is a BaseException, so a task cancelled during the first await in that block (release_page_with_context) propagates out and page.close() is never reached — the page and its context refcount leak. Reachable from dispatcher/stream teardown paths that cancel in-flight tasks.
Impact
In long-lived processes (the Docker server's browser pool, any SDK user reusing a crawler), leaked pages accumulate indefinitely. Combined with the launch flags that disable background throttling, each leaked page also burns CPU forever. This is the core mechanism behind the multi-week renderer accumulation in #2202.
Proposed fix
- Wrap the post-navigation phase (or at minimum every
evaluate/content()call) inasyncio.wait_forwith a budget derived frompage_timeout, so a crawl always terminates. - Make the cleanup
finallycancellation-safe: catchBaseException(re-raisingCancelledErrorafter the page is closed) or shield the close.
- Lenguaje dominante
- Python
- Estrellas
- 84.5k
- Forks
- 8.7k
- Merge medio
- 3 d 20 h
- PR fusionados (30 d)
- 15
Preparar el entorno
- Incluye un Dockerfile o un archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de unclecode/crawl4ai
-
🐞 Bug 🩺 Needs Triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
unclecode/crawl4ai#2319 · 2 comentarios ·
Los mantenedores suelen responder en 1 día
-
[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh runAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
unclecode/crawl4ai#2309 · 2 comentarios ·
Los mantenedores suelen responder en 1 día
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 84/100
unclecode/crawl4ai#2147 · 3 comentarios ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
unclecode/crawl4ai#2123 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
🐞 Bug 🩺 Needs Triage
Dificultad 4/5 3-5 días Aptitud para principiantes 55/100
Los mantenedores suelen responder en 1 día
Todos los issues de unclecode/crawl4ai
Issues similares
-
repo-audit
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
scverse/repo-health#20 ·
Los mantenedores suelen responder en 1 día
-
/context/prime scope override double-prefixes an entity-ref project and drops its scoped memoriesAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
phasespace-labs/palinode#232 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
collective/icalendar#1858 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
Los mantenedores suelen responder en 1 día
-
lfx-mcp cannot supply global variables: LangflowClient drops X-LANGFLOW-GLOBAL-VAR-* from envAbiertobug
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
langflow-ai/langflow#15496 ·
Los mantenedores suelen responder en 1 día