Feature request: opt-in block_internal_urls egress filter for the library layer
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 45/100
- Loại issue
- Tính năng
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Ít trao đổi
- Công nghệ
- docker, playwright, python
- Lĩnh vực
- backend, networking, security
Hướng nghiên cứu
Read crawl4ai/async_crawler_strategy.py and trace the HTTP path through AsyncHTTPCrawlerStrategy._handle_http() and the browser path through browser_manager.py to Playwright goto(). Then compare validate_url_destination and resolve_and_pin in deploy/docker/utils.py and egress_broker.py; the work is done when one shared opt-in filter covers DNS resolution, redirects, and both egress paths without fetching blocked destinations.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Thanks for following up from the email thread, @ntohidi.
Quick correction on my side: the issue body initially only contained a literal file path — I mistakenly relied on @path expansion with gh api, which doesn't expand files (that's a curl / --body-file flag). Pasting the real content here.
Alignment
We're aligned on the classification. The library is a user agent invoked by a trusted caller, so destination filtering should stay opt-in, and the SSRF trust boundary remains at the Docker API server where egress_broker.py already enforces it. The agentic / LLM-chosen-URL case is the scenario that justifies exposing the same primitives to library callers.
Proposed design
- Flag:
block_internal_urls: bool(defaultFalse, opt-in). Set per-crawl so callers who embed Crawl4AI in an agent can opt in without a global change. - Chokepoint: a single host-validation call inserted right after the existing scheme allow-list check in
AsyncCrawlerStrategy.crawl()(crawl4ai/async_crawler_strategy.py). It must cover both egress paths:- HTTP path:
AsyncHTTPCrawlerStrategy._handle_http()(aiohttp) - Browser path:
browser_manager.py→ Playwrightgoto()
- HTTP path:
- Logic reuse: port the existing
validate_url_destination+resolve_and_pin(DNS pinning) + per-hop redirect revalidation fromdeploy/docker/utils.py/egress_broker.pyinto a shared helper (e.g.crawl4ai/url_safety.py) so the library and the Docker server share one implementation — no duplicated trust logic. - Blocked ranges: loopback (
127.0.0.0/8,::1), private (10/8,172.16/12,192.168/16,fc00::/7), link-local (169.254/16incl. cloud metadata169.254.169.254,fe80::/10), and0.0.0.0/8. The resolved IP must be checked after DNS (pin) and after every redirect hop (revalidate) to prevent DNS-rebinding / redirect-to-internal bypasses. - Behavior on block: raise a
BlockedURLexception (or return a failedCrawlResultwith a clear error) rather than fetching.
Open questions (happy to match maintainer preference)
- Flag name —
block_internal_urlsvsdeny_private_destinationsvsegress_filter? - Where the chokepoint sits — shared base
crawl()vs per-strategy hooks? - Browser redirect following — should the redirect revalidation also cover hops taken by Playwright
goto, or only the initial URL?
I'd be happy to draft the PR implementing this once we settle the surface.
- Ngôn ngữ chính
- Python
- Star
- 84.5k
- Fork
- 8.7k
- Merge trung bình
- 3 ngày 20 giờ
- Pull request đã merge (30 ngày)
- 15
Chuẩn bị môi trường
- Có Dockerfile hoặc tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của unclecode/crawl4ai
-
🐞 Bug 🩺 Needs Triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
unclecode/crawl4ai#2319 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh runĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
unclecode/crawl4ai#2309 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 84/100
unclecode/crawl4ai#2147 · 3 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
unclecode/crawl4ai#2123 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
🐞 Bug 🩺 Needs Triage
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 55/100
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của unclecode/crawl4ai
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
kornia/kornia#5263 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Metadata correction for W16-5400Đang mởapproved correction metadata
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
acl-org/acl-anthology#10133 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
BasedHardware/omi#20084 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug needs-acceptance wg/evaluation-quality
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
vllm-project/semantic-router#4424 ·
Maintainer thường phản hồi trong vòng 1 ngày