[Bug]: BFSDeepCrawlStrategy.can_process_url() rejects valid single-label hostnames (netloc without a dot)
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 3/5
- Thời gian dự kiến
- 1-2 ngày
- Mức phù hợp với người mới
- 68/100
Hướng nghiên cứu
Start by locating BFSDeepCrawlStrategy.can_process_url() and trace how discovered links reach filter_chain.apply(); reproduce the failure with https://name/xyz/ and a single-label hostname. Done means valid http/https URLs with a non-empty netloc are accepted and the configured DomainFilter is reached for links beyond depth 0.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
crawl4ai version
0.9.2
Expected Behavior
Single-label/internal hostnames without a dot should be treated as valid
hostnames, at least when the URL has a valid scheme (http/https) and
non-empty netloc. The dot-check appears to be an overly strict heuristic
that isn't part of the documented API/config (no allow_single_label_hosts
or similar flag exists).
Current Behavior
- Depth 0 (start URL) is fetched successfully.
- Every discovered link at depth > 0 is logged as:
Invalid URL: https://name/xyz/xyz, error: Invalid domain
The configuredDomainFilteris never even reached, since the exception
is raised beforefilter_chain.apply()is called.
Is this reproducible?
Yes
Inputs Causing the Bug
Steps to Reproduce
1. Start URL: `https://name/xyz/` (a host with no dot in its name,
reachable e.g. via internal DNS or /etc/hosts).
2. Configure `BFSDeepCrawlStrategy` with
`filter_chain=FilterChain([DomainFilter(allowed_domains=["iltis3"])])`.
3. Run `crawler.arun()`.
Code snippets
OS
Linux
Python version
3.12.13
Browser
No response
Browser version
No response
Error logs & Screenshots (if applicable)
[INIT].... → Crawl4AI 0.9.2
[FETCH]... ↓ https://name/xyz | ✓ | ⏱: 0.91s
[SCRAPE].. ◆ https://name/xyz | ✓ | ⏱: 0.06s
[EXTRACT]. ■ https://name/xyz | ✓ | ⏱: 0.03s
[COMPLETE] ● https://name/xyz | ✓ | ⏱: 1.01s
Invalid URL: https://name/xyz/neu-im-intranet, error: Invalid domain
Invalid URL: name/xyz/suche, error: Invalid domain
Invalid URL: https://name/xyz
/allgemein/stellen, error: Invalid domain
- Ngôn ngữ chính
- Python
- Star
- 84.5k
- Fork
- 8.7k
- Merge trung bình
- 3 ngày 9 giờ
- Pull request đã merge (30 ngày)
- 17
Chuẩn bị môi trường
- Có Dockerfile hoặc tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của unclecode/crawl4ai
-
[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh runĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 84/100
unclecode/crawl4ai#2147 · 3 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
unclecode/crawl4ai#2123 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
🐞 Bug 🩺 Needs Triage
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 55/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của unclecode/crawl4ai
Issue tương tự
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 90/100
Maintainer thường phản hồi trong vòng 1 ngày
-
https://search.utilibre.orgĐang mởinstance instance add
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
searxng/searx-instances#941 · 1 bình luận ·
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
FluidNumerics/fluid-walk-blocker#89 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
Maintainer thường phản hồi trong vòng 1 ngày