[Bug]: One seed that fails the destination check rejects the whole `/crawl` batch
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 3/5
- 見積もり時間
- 1〜2日
- 初心者へのやさしさ
- 72/100
調査の方向性
Start at _normalize_and_validate_seeds in deploy/docker/api.py around lines 650-657 and trace how /crawl collects results. Then inspect resolve_and_pin in deploy/docker/egress_broker.py around lines 92-96, reproducing the two-seed request from the issue. Done means one refused or unresolvable seed becomes an opaque failed result while other seeds still crawl successfully.
索引モデルが issue の本文から書いたものです。
説明
crawl4ai version
0.9.4
Expected Behavior
The other seeds are crawled, and the refused one comes back as a failed result carrying the same opaque message, the way a robots.txt refusal already does inside a batch.
Current Behavior
_normalize_and_validate_seeds checks every seed before crawling and raises on the first one refused (api.py#L650-L657), so the whole request fails with 400 {"detail": "URL blocked (SSRF protection): URL blocked"} and no results. resolve_and_pin also raises EgressBlocked when getaddrinfo fails (egress_broker.py#L92-L96), so a host that simply doesn't resolve (a dead domain from a stale sitemap) takes the batch down the same way. The caller can't tell which seed it was, so its only recovery is to split the batch and retry.
Keeping the message opaque, with no resolution oracle, is right, and a per-URL failure keeps it: it says no more than today's 400, because a caller can already find the refused seed by sending the URLs one at a time.
If per-URL failures are the direction you want, I can send a PR. cc @ntohidi
Is this reproducible?
Yes
Inputs Causing the Bug
- urls: ["https://example.com/", "https://no-such-host-12345.example/"]
Steps to Reproduce
1. Run the Docker image (0.9.4, or built from develop), without CRAWL4AI_ALLOW_INTERNAL_URLS
2. POST /crawl with the two URLs above
3. 400, no results; the first URL alone crawls fine
Code snippets
import httpx
r = httpx.post("http://localhost:11235/crawl", headers={"Authorization": f"Bearer {TOKEN}"},
json={"urls": ["https://example.com/", "https://no-such-host-12345.example/"]}, timeout=60)
print(r.status_code, r.text) # 400 {"detail":"URL blocked (SSRF protection): URL blocked"}
OS
Linux (Docker image built from develop @ 1f68e5b)
Python version
3.12 (the image's)
Browser
Chromium (Playwright, bundled)
Browser version
No response
Error logs & Screenshots (if applicable)
[seedbatch] one resolvable + one unresolvable seed -> HTTP 400: {"detail":"URL blocked (SSRF protection): URL blocked"}
[seedbatch] the resolvable seed alone -> HTTP 200, success=True
- 主要言語
- Python
- スター
- 84.5k
- フォーク
- 8.7k
- 平均マージ
- 3日 9時間
- マージ済み PR(30日)
- 17
環境構築
- Dockerfile または Docker Compose ファイルあり
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
unclecode/crawl4ai のほかの issue
-
[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh runオープン
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
unclecode/crawl4ai#2309 · コメント 2 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 84/100
unclecode/crawl4ai#2147 · コメント 3 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
unclecode/crawl4ai#2123 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
🐞 Bug 🩺 Needs Triage
難易度 4/5 3〜5日 初心者へのやさしさ 55/100
メンテナーはふだん 1 日以内に返信
-
難易度 5/5 1週間以上 初心者へのやさしさ 35/100
メンテナーはふだん 1 日以内に返信
unclecode/crawl4ai の issue をすべて見る
似ている issue
-
難易度 1/5 1時間未満 初心者へのやさしさ 72/100
letsencrypt/cp-cps#353 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
-
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
PedestrianDynamics/pyFDS-Evac#394 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
DOI-USGS/pywatershed#421 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
python-pillow/Pillow#10087 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信