Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh run

オープン 初心者向け
#2,309 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
2/5
見積もり時間
1〜3時間
初心者へのやさしさ
78/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
python
領域
backend

調査の方向性

BFSDeepCrawlStrategy から始め、_arun_batch と _arun_stream の非再開初期化分岐を調査します。リンクされた回帰ハーネス内の決定論的なクローラーを使って、batch、stream、混在モードでの繰り返し呼び出しを検証し、再開されたクロールでは保存済みのカウンターを維持します。12 ケースのハーネスがパスし、新しい実行ごとに独自の max_pages バジェットが与えられれば完了です。

索引モデルが issue の本文から書いたものです。

説明

Description

BFSDeepCrawlStrategy resets its cancellation event at the start of each batch/stream run, but the fresh-run branch does not reset _pages_crawled. Consequently, sequential independent crawls using the same strategy consume one shared page budget. A later run can return no results even though its start URL was never crawled.

This is distinct from resuming a saved crawl: _resume_state is absent in the failing case. A resumed run should continue to use its saved counter.

Minimal reproduction

This uses the public strategy with a deterministic crawler boundary; no browser or HTTP request is needed:

import asyncio
from types import SimpleNamespace
from crawl4ai import CrawlerRunConfig
from crawl4ai.deep_crawling import BFSDeepCrawlStrategy

class Crawler:
    async def arun_many(self, urls, config):
        return [SimpleNamespace(url=u, success=True, metadata={}, links={'internal': []}) for u in urls]

async def main():
    strategy = BFSDeepCrawlStrategy(max_depth=0, max_pages=1)
    config = CrawlerRunConfig(stream=False)
    for url in ['https://example.test/first', 'https://example.test/second']:
        results = await strategy.arun(start_url=url, crawler=Crawler(), config=config)
        print([r.url for r in results])

asyncio.run(main())

Expected: each call returns its respective start URL. Current behavior: the first call returns its URL; the second returns [] before invoking the crawler.

Candidate fix

Reset self._pages_crawled = 0 in the non-resume initialization branch of both _arun_batch and _arun_stream. Leave the resume branch's pages_crawled restoration unchanged.

Validation

Tested develop at 1f68e5bd7c29f2067a1ef74f28dbf4dc20686a06.

Before: 10 failed, 2 passed. After: 12 passed. Tests exercise the complete BFS/base strategy and filter/scorer modules with a deterministic async crawler and lightweight configuration/result/statistics objects. They cover repeated batch/stream/mixed-mode calls, cancellation followed by a fresh run, and preservation of a saved resume counter. All test crawls have max_depth=0, so URL normalization and browser/network behavior are outside this test scope. The full repository suite was not run.

Python 3.13.5, pytest 9.0.2, Linux. Issue searches for reuse max_pages and _pages_crawled reset, and PR search for the latter, did not find a specific existing report of sequential fresh-run counter leakage.

Patch and tests

Source-pinned patch and 12-case regression harness. From the bundle root, run python reproduce.py --fetch --case crawl4ai-fresh-run-budget --variant after. The runner checks the original source hash and patch application before testing.

主要言語
Python
スター
84.5k
フォーク
8.7k
平均マージ
3日 9時間
マージ済み PR(30日)
17

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

unclecode/crawl4ai のほかの issue

unclecode/crawl4ai の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。