[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh run
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
調査の方向性
BFSDeepCrawlStrategy から始め、_arun_batch と _arun_stream の非再開初期化分岐を調査します。リンクされた回帰ハーネス内の決定論的なクローラーを使って、batch、stream、混在モードでの繰り返し呼び出しを検証し、再開されたクロールでは保存済みのカウンターを維持します。12 ケースのハーネスがパスし、新しい実行ごとに独自の max_pages バジェットが与えられれば完了です。
索引モデルが issue の本文から書いたものです。
説明
Description
BFSDeepCrawlStrategy resets its cancellation event at the start of each batch/stream run, but the fresh-run branch does not reset _pages_crawled. Consequently, sequential independent crawls using the same strategy consume one shared page budget. A later run can return no results even though its start URL was never crawled.
This is distinct from resuming a saved crawl: _resume_state is absent in the failing case. A resumed run should continue to use its saved counter.
Minimal reproduction
This uses the public strategy with a deterministic crawler boundary; no browser or HTTP request is needed:
import asyncio
from types import SimpleNamespace
from crawl4ai import CrawlerRunConfig
from crawl4ai.deep_crawling import BFSDeepCrawlStrategy
class Crawler:
async def arun_many(self, urls, config):
return [SimpleNamespace(url=u, success=True, metadata={}, links={'internal': []}) for u in urls]
async def main():
strategy = BFSDeepCrawlStrategy(max_depth=0, max_pages=1)
config = CrawlerRunConfig(stream=False)
for url in ['https://example.test/first', 'https://example.test/second']:
results = await strategy.arun(start_url=url, crawler=Crawler(), config=config)
print([r.url for r in results])
asyncio.run(main())
Expected: each call returns its respective start URL. Current behavior: the first call returns its URL; the second returns [] before invoking the crawler.
Candidate fix
Reset self._pages_crawled = 0 in the non-resume initialization branch of both _arun_batch and _arun_stream. Leave the resume branch's pages_crawled restoration unchanged.
Validation
Tested develop at 1f68e5bd7c29f2067a1ef74f28dbf4dc20686a06.
Before: 10 failed, 2 passed. After: 12 passed. Tests exercise the complete BFS/base strategy and filter/scorer modules with a deterministic async crawler and lightweight configuration/result/statistics objects. They cover repeated batch/stream/mixed-mode calls, cancellation followed by a fresh run, and preservation of a saved resume counter. All test crawls have max_depth=0, so URL normalization and browser/network behavior are outside this test scope. The full repository suite was not run.
Python 3.13.5, pytest 9.0.2, Linux. Issue searches for reuse max_pages and _pages_crawled reset, and PR search for the latter, did not find a specific existing report of sequential fresh-run counter leakage.
Patch and tests
Source-pinned patch and 12-case regression harness. From the bundle root, run python reproduce.py --fetch --case crawl4ai-fresh-run-budget --variant after. The runner checks the original source hash and patch application before testing.
- 主要言語
- Python
- スター
- 84.5k
- フォーク
- 8.7k
- 平均マージ
- 3日 9時間
- マージ済み PR(30日)
- 17
環境構築
- Dockerfile または Docker Compose ファイルあり
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
unclecode/crawl4ai のほかの issue
-
難易度 1/5 1時間未満 初心者へのやさしさ 84/100
unclecode/crawl4ai#2147 · コメント 3 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
unclecode/crawl4ai#2123 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
🐞 Bug 🩺 Needs Triage
難易度 4/5 3〜5日 初心者へのやさしさ 55/100
メンテナーはふだん 1 日以内に返信
-
難易度 5/5 1週間以上 初心者へのやさしさ 35/100
メンテナーはふだん 1 日以内に返信
-
難易度 3/5 1〜2日 初心者へのやさしさ 58/100
メンテナーはふだん 1 日以内に返信
unclecode/crawl4ai の issue をすべて見る
似ている issue
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 90/100
メンテナーはふだん 1 日以内に返信
-
instance instance add
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
searxng/searx-instances#941 · コメント 1 件 ·
-
難易度 1/5 1時間未満 初心者へのやさしさ 92/100
FluidNumerics/fluid-walk-blocker#89 ·
メンテナーはふだん 1 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
メンテナーはふだん 1 日以内に返信