Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Feature request: opt-in block_internal_urls egress filter for the library layer

オープン
#2,146 コメント 2 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
45/100
issue の種類
機能追加
明瞭さ
おおむね明確
活発さ
静か
技術スタック
docker, playwright, python

調査の方向性

Read crawl4ai/async_crawler_strategy.py and trace the HTTP path through AsyncHTTPCrawlerStrategy._handle_http() and the browser path through browser_manager.py to Playwright goto(). Then compare validate_url_destination and resolve_and_pin in deploy/docker/utils.py and egress_broker.py; the work is done when one shared opt-in filter covers DNS resolution, redirects, and both egress paths without fetching blocked destinations.

索引モデルが issue の本文から書いたものです。

説明

✨ Enhancement

Thanks for following up from the email thread, @ntohidi.

Quick correction on my side: the issue body initially only contained a literal file path — I mistakenly relied on @path expansion with gh api, which doesn't expand files (that's a curl / --body-file flag). Pasting the real content here.

Alignment

We're aligned on the classification. The library is a user agent invoked by a trusted caller, so destination filtering should stay opt-in, and the SSRF trust boundary remains at the Docker API server where egress_broker.py already enforces it. The agentic / LLM-chosen-URL case is the scenario that justifies exposing the same primitives to library callers.

Proposed design
  • Flag: block_internal_urls: bool (default False, opt-in). Set per-crawl so callers who embed Crawl4AI in an agent can opt in without a global change.
  • Chokepoint: a single host-validation call inserted right after the existing scheme allow-list check in AsyncCrawlerStrategy.crawl() (crawl4ai/async_crawler_strategy.py). It must cover both egress paths:
    • HTTP path: AsyncHTTPCrawlerStrategy._handle_http() (aiohttp)
    • Browser path: browser_manager.py → Playwright goto()
  • Logic reuse: port the existing validate_url_destination + resolve_and_pin (DNS pinning) + per-hop redirect revalidation from deploy/docker/utils.py / egress_broker.py into a shared helper (e.g. crawl4ai/url_safety.py) so the library and the Docker server share one implementation — no duplicated trust logic.
  • Blocked ranges: loopback (127.0.0.0/8, ::1), private (10/8, 172.16/12, 192.168/16, fc00::/7), link-local (169.254/16 incl. cloud metadata 169.254.169.254, fe80::/10), and 0.0.0.0/8. The resolved IP must be checked after DNS (pin) and after every redirect hop (revalidate) to prevent DNS-rebinding / redirect-to-internal bypasses.
  • Behavior on block: raise a BlockedURL exception (or return a failed CrawlResult with a clear error) rather than fetching.
Open questions (happy to match maintainer preference)
  1. Flag name — block_internal_urls vs deny_private_destinations vs egress_filter?
  2. Where the chokepoint sits — shared base crawl() vs per-strategy hooks?
  3. Browser redirect following — should the redirect revalidation also cover hops taken by Playwright goto, or only the initial URL?

I'd be happy to draft the PR implementing this once we settle the surface.

主要言語
Python
スター
84.5k
フォーク
8.7k
平均マージ
3日 20時間
マージ済み PR(30日)
15

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

unclecode/crawl4ai のほかの issue

unclecode/crawl4ai の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。