Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh run

Open Beginner friendly
#2,309 2 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
78/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python
Domain
backend

Research direction

Start with BFSDeepCrawlStrategy and inspect the non-resume initialization branches of _arun_batch and _arun_stream. Verify repeated batch, stream, and mixed-mode calls with the deterministic crawler in the linked regression harness, while preserving the saved counter for resumed crawls. Done means the 12-case harness passes and each fresh run receives its own max_pages budget.

Written by the indexing model from the issue text.

Description

Description

BFSDeepCrawlStrategy resets its cancellation event at the start of each batch/stream run, but the fresh-run branch does not reset _pages_crawled. Consequently, sequential independent crawls using the same strategy consume one shared page budget. A later run can return no results even though its start URL was never crawled.

This is distinct from resuming a saved crawl: _resume_state is absent in the failing case. A resumed run should continue to use its saved counter.

Minimal reproduction

This uses the public strategy with a deterministic crawler boundary; no browser or HTTP request is needed:

import asyncio
from types import SimpleNamespace
from crawl4ai import CrawlerRunConfig
from crawl4ai.deep_crawling import BFSDeepCrawlStrategy

class Crawler:
    async def arun_many(self, urls, config):
        return [SimpleNamespace(url=u, success=True, metadata={}, links={'internal': []}) for u in urls]

async def main():
    strategy = BFSDeepCrawlStrategy(max_depth=0, max_pages=1)
    config = CrawlerRunConfig(stream=False)
    for url in ['https://example.test/first', 'https://example.test/second']:
        results = await strategy.arun(start_url=url, crawler=Crawler(), config=config)
        print([r.url for r in results])

asyncio.run(main())

Expected: each call returns its respective start URL. Current behavior: the first call returns its URL; the second returns [] before invoking the crawler.

Candidate fix

Reset self._pages_crawled = 0 in the non-resume initialization branch of both _arun_batch and _arun_stream. Leave the resume branch's pages_crawled restoration unchanged.

Validation

Tested develop at 1f68e5bd7c29f2067a1ef74f28dbf4dc20686a06.

Before: 10 failed, 2 passed. After: 12 passed. Tests exercise the complete BFS/base strategy and filter/scorer modules with a deterministic async crawler and lightweight configuration/result/statistics objects. They cover repeated batch/stream/mixed-mode calls, cancellation followed by a fresh run, and preservation of a saved resume counter. All test crawls have max_depth=0, so URL normalization and browser/network behavior are outside this test scope. The full repository suite was not run.

Python 3.13.5, pytest 9.0.2, Linux. Issue searches for reuse max_pages and _pages_crawled reset, and PR search for the latter, did not find a specific existing report of sequential fresh-run counter leakage.

Patch and tests

Source-pinned patch and 12-case regression harness. From the bundle root, run python reproduce.py --fetch --case crawl4ai-fresh-run-budget --variant after. The runner checks the original source hash and patch application before testing.

Dominant language
Python
Stars
84.5k
Forks
8.7k
Avg merge
3d 9h
Merged PRs (30d)
17

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from unclecode/crawl4ai

All issues in unclecode/crawl4ai

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.