[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh run
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 2/5
- Tempo stimato
- 1-3 ore
- Idoneità per principianti
- 78/100
Direzione di ricerca
Inizia con BFSDeepCrawlStrategy e ispeziona i rami di inizializzazione non-resume di _arun_batch e _arun_stream. Verifica chiamate ripetute in modalità batch, stream e mista con il crawler deterministico nell’harness di regressione collegato, preservando il contatore salvato per i crawl ripresi. Il lavoro è completato quando l’harness a 12 casi passa e ogni nuova esecuzione riceve il proprio budget max_pages.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Description
BFSDeepCrawlStrategy resets its cancellation event at the start of each batch/stream run, but the fresh-run branch does not reset _pages_crawled. Consequently, sequential independent crawls using the same strategy consume one shared page budget. A later run can return no results even though its start URL was never crawled.
This is distinct from resuming a saved crawl: _resume_state is absent in the failing case. A resumed run should continue to use its saved counter.
Minimal reproduction
This uses the public strategy with a deterministic crawler boundary; no browser or HTTP request is needed:
import asyncio
from types import SimpleNamespace
from crawl4ai import CrawlerRunConfig
from crawl4ai.deep_crawling import BFSDeepCrawlStrategy
class Crawler:
async def arun_many(self, urls, config):
return [SimpleNamespace(url=u, success=True, metadata={}, links={'internal': []}) for u in urls]
async def main():
strategy = BFSDeepCrawlStrategy(max_depth=0, max_pages=1)
config = CrawlerRunConfig(stream=False)
for url in ['https://example.test/first', 'https://example.test/second']:
results = await strategy.arun(start_url=url, crawler=Crawler(), config=config)
print([r.url for r in results])
asyncio.run(main())
Expected: each call returns its respective start URL. Current behavior: the first call returns its URL; the second returns [] before invoking the crawler.
Candidate fix
Reset self._pages_crawled = 0 in the non-resume initialization branch of both _arun_batch and _arun_stream. Leave the resume branch's pages_crawled restoration unchanged.
Validation
Tested develop at 1f68e5bd7c29f2067a1ef74f28dbf4dc20686a06.
Before: 10 failed, 2 passed. After: 12 passed. Tests exercise the complete BFS/base strategy and filter/scorer modules with a deterministic async crawler and lightweight configuration/result/statistics objects. They cover repeated batch/stream/mixed-mode calls, cancellation followed by a fresh run, and preservation of a saved resume counter. All test crawls have max_depth=0, so URL normalization and browser/network behavior are outside this test scope. The full repository suite was not run.
Python 3.13.5, pytest 9.0.2, Linux. Issue searches for reuse max_pages and _pages_crawled reset, and PR search for the latter, did not find a specific existing report of sequential fresh-run counter leakage.
Patch and tests
Source-pinned patch and 12-case regression harness. From the bundle root, run python reproduce.py --fetch --case crawl4ai-fresh-run-budget --variant after. The runner checks the original source hash and patch application before testing.
- Lingua principale
- Python
- Stelle
- 84.5k
- Fork
- 8.7k
- Merge medio
- 3g 9h
- PR unite (30g)
- 17
Preparare l'ambiente
- Include un Dockerfile o un file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di unclecode/crawl4ai
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 84/100
unclecode/crawl4ai#2147 · 3 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
unclecode/crawl4ai#2123 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
🐞 Bug 🩺 Needs Triage
Difficoltà 4/5 3-5 giorni Idoneità per principianti 55/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 58/100
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di unclecode/crawl4ai
Issue simili
-
[Bug] @deck.gl/arcgis dist import resolves to unpublished @deck.gl/core source path (9.3.11, 9.4.0)Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno
-
workflow: a tick's dispatch counts as 'only this step', and no review self-grants a round unattendedApertaworkflow
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
kristofdegrave/homeassistant-smart-charging#1505 ·
I maintainer di solito rispondono entro 1 giorno
-
New Submission: TropWATERApertametadata submission
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
canonical/content-cache-operator#163 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
[submission]Apertasubmission
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 65/100
leanprover/lean-eval-submissions#1852 ·
I maintainer di solito rispondono entro 1 giorno