Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh run

未关闭 适合新手
#2,309 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
2/5
预计耗时
1-3 小时
新手友好度
78/100
Issue 类型
缺陷
描述清晰度
描述清楚
活跃度
活跃
技术栈
python
领域
backend

调研方向

从 BFSDeepCrawlStrategy 开始,检查 _arun_batch 和 _arun_stream 的非恢复初始化分支。使用链接的回归测试 harness 中的确定性爬虫,验证 batch、stream 和混合模式下的重复调用,同时保留恢复爬取所使用的已保存计数器。完成的标准是 12 个用例的 harness 通过,并且每次全新运行都获得自己的 max_pages 预算。

由索引模型根据 Issue 内容生成。

描述

Description

BFSDeepCrawlStrategy resets its cancellation event at the start of each batch/stream run, but the fresh-run branch does not reset _pages_crawled. Consequently, sequential independent crawls using the same strategy consume one shared page budget. A later run can return no results even though its start URL was never crawled.

This is distinct from resuming a saved crawl: _resume_state is absent in the failing case. A resumed run should continue to use its saved counter.

Minimal reproduction

This uses the public strategy with a deterministic crawler boundary; no browser or HTTP request is needed:

import asyncio
from types import SimpleNamespace
from crawl4ai import CrawlerRunConfig
from crawl4ai.deep_crawling import BFSDeepCrawlStrategy

class Crawler:
    async def arun_many(self, urls, config):
        return [SimpleNamespace(url=u, success=True, metadata={}, links={'internal': []}) for u in urls]

async def main():
    strategy = BFSDeepCrawlStrategy(max_depth=0, max_pages=1)
    config = CrawlerRunConfig(stream=False)
    for url in ['https://example.test/first', 'https://example.test/second']:
        results = await strategy.arun(start_url=url, crawler=Crawler(), config=config)
        print([r.url for r in results])

asyncio.run(main())

Expected: each call returns its respective start URL. Current behavior: the first call returns its URL; the second returns [] before invoking the crawler.

Candidate fix

Reset self._pages_crawled = 0 in the non-resume initialization branch of both _arun_batch and _arun_stream. Leave the resume branch's pages_crawled restoration unchanged.

Validation

Tested develop at 1f68e5bd7c29f2067a1ef74f28dbf4dc20686a06.

Before: 10 failed, 2 passed. After: 12 passed. Tests exercise the complete BFS/base strategy and filter/scorer modules with a deterministic async crawler and lightweight configuration/result/statistics objects. They cover repeated batch/stream/mixed-mode calls, cancellation followed by a fresh run, and preservation of a saved resume counter. All test crawls have max_depth=0, so URL normalization and browser/network behavior are outside this test scope. The full repository suite was not run.

Python 3.13.5, pytest 9.0.2, Linux. Issue searches for reuse max_pages and _pages_crawled reset, and PR search for the latter, did not find a specific existing report of sequential fresh-run counter leakage.

Patch and tests

Source-pinned patch and 12-case regression harness. From the bundle root, run python reproduce.py --fetch --case crawl4ai-fresh-run-budget --variant after. The runner checks the original source hash and patch application before testing.

主要语言
Python
星标
84.5k
派生
8.7k
平均合并
3 天 9 小时
30 天内合并 PR
17

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

unclecode/crawl4ai 的其他 Issue

查看 unclecode/crawl4ai 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。