Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[Bug]: Docker pool's permanent browser never serves any request - signature is computed without the egress proxy, so no `/crawl` can ever match it

クローズ
#2,204 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
3/5
見積もり時間
1〜2日
初心者へのやさしさ
74/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
python

調査の方向性

Start with crawler_pool.py:46-49, 60, 177, and 198, then trace the permanent-browser setup in server.py:126-141 and 199-203 and request handling in api.py:687-690. Verify that the permanent browser uses a matching request signature, serves the default traffic, and is included in janitor cleanup without creating a second browser tree.

索引モデルが issue の本文から書いたものです。

説明

⚙ Done 🐞 Bug

Summary

Found while investigating #2202. The Docker server's "permanent" warm browser is dead weight: it is started at boot, serves zero requests, is never cleaned up, and causes every container to run a second browser tree from the first request onward.

Root cause

The pool matches requests to browsers by hashing the full BrowserConfig (crawler_pool.py:46-49), and proxy_config is part of to_dict() (async_configs.py:968).

  • init_permanent() is called at startup with a config built straight from config.yml — without enforce_egress() (server.py:199-203), so its fingerprint has proxy_config: None.
  • Every /crawl request's config goes through enforce_egress() first (api.py:687-690), which sets proxy_config to the egress pinning proxy — whose port is random each boot (egress_broker.py:198-201).

The two fingerprints can never be equal, so _is_default_config() (crawler_pool.py:60) never matches and PERMANENT is never returned. Note get_default_browser_config() (server.py:126-141) does apply enforce_egress, so even the server's own /html, /screenshot, /pdf, /execute_js endpoints miss it.

Additionally, the janitor only sweeps HOT_POOL and COLD_POOL (crawler_pool.py:177,198) — PERMANENT is never inspected, so the unused browser also can never be reclaimed.

Impact

  • ~270 MB RSS + one Playwright driver + full Chromium tree per container, doing nothing, forever.
  • Every container shows two driver → browser process branches after the first request (observed in #2202's process listing: one branch from boot, one from the first crawl a day later).
  • The text_mode optimization configured for the default browser never applies to real traffic.

Verified

Reproduced on unclecode/crawl4ai:0.9.2: fresh container = one browser tree; after a single plain /crawl = two independent trees, permanent one idle.

Proposed fix

Keep the permanent browser, fix the match: build its config through the same path requests use (get_default_browser_config(), i.e. after enforce_egress), and compute DEFAULT_CONFIG_SIG from that. Alternatively/additionally, exclude the server-injected proxy_config from the pool signature, since the server sets it identically on every request.

主要言語
Python
スター
84.5k
フォーク
8.7k
平均マージ
3日 9時間
マージ済み PR(30日)
17

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

unclecode/crawl4ai のほかの issue

unclecode/crawl4ai の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。