Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh run

Aperta Adatta ai principianti
#2,309 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
2/5
Tempo stimato
1-3 ore
Idoneità per principianti
78/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
python
Ambito
backend

Direzione di ricerca

Inizia con BFSDeepCrawlStrategy e ispeziona i rami di inizializzazione non-resume di _arun_batch e _arun_stream. Verifica chiamate ripetute in modalità batch, stream e mista con il crawler deterministico nell’harness di regressione collegato, preservando il contatore salvato per i crawl ripresi. Il lavoro è completato quando l’harness a 12 casi passa e ogni nuova esecuzione riceve il proprio budget max_pages.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Description

BFSDeepCrawlStrategy resets its cancellation event at the start of each batch/stream run, but the fresh-run branch does not reset _pages_crawled. Consequently, sequential independent crawls using the same strategy consume one shared page budget. A later run can return no results even though its start URL was never crawled.

This is distinct from resuming a saved crawl: _resume_state is absent in the failing case. A resumed run should continue to use its saved counter.

Minimal reproduction

This uses the public strategy with a deterministic crawler boundary; no browser or HTTP request is needed:

import asyncio
from types import SimpleNamespace
from crawl4ai import CrawlerRunConfig
from crawl4ai.deep_crawling import BFSDeepCrawlStrategy

class Crawler:
    async def arun_many(self, urls, config):
        return [SimpleNamespace(url=u, success=True, metadata={}, links={'internal': []}) for u in urls]

async def main():
    strategy = BFSDeepCrawlStrategy(max_depth=0, max_pages=1)
    config = CrawlerRunConfig(stream=False)
    for url in ['https://example.test/first', 'https://example.test/second']:
        results = await strategy.arun(start_url=url, crawler=Crawler(), config=config)
        print([r.url for r in results])

asyncio.run(main())

Expected: each call returns its respective start URL. Current behavior: the first call returns its URL; the second returns [] before invoking the crawler.

Candidate fix

Reset self._pages_crawled = 0 in the non-resume initialization branch of both _arun_batch and _arun_stream. Leave the resume branch's pages_crawled restoration unchanged.

Validation

Tested develop at 1f68e5bd7c29f2067a1ef74f28dbf4dc20686a06.

Before: 10 failed, 2 passed. After: 12 passed. Tests exercise the complete BFS/base strategy and filter/scorer modules with a deterministic async crawler and lightweight configuration/result/statistics objects. They cover repeated batch/stream/mixed-mode calls, cancellation followed by a fresh run, and preservation of a saved resume counter. All test crawls have max_depth=0, so URL normalization and browser/network behavior are outside this test scope. The full repository suite was not run.

Python 3.13.5, pytest 9.0.2, Linux. Issue searches for reuse max_pages and _pages_crawled reset, and PR search for the latter, did not find a specific existing report of sequential fresh-run counter leakage.

Patch and tests

Source-pinned patch and 12-case regression harness. From the bundle root, run python reproduce.py --fetch --case crawl4ai-fresh-run-budget --variant after. The runner checks the original source hash and patch application before testing.

Lingua principale
Python
Stelle
84.5k
Fork
8.7k
Merge medio
3g 9h
PR unite (30g)
17

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di unclecode/crawl4ai

Tutte le issue di unclecode/crawl4ai

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.