Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

[Bug]: BFSDeepCrawlStrategy.can_process_url() rejects valid single-label hostnames (netloc without a dot)

Đang mở
#2,079 1 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
3/5
Thời gian dự kiến
1-2 ngày
Mức phù hợp với người mới
68/100
Loại issue
Lỗi
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Ít trao đổi
Công nghệ
python
Lĩnh vực
web-dev

Hướng nghiên cứu

Start by locating BFSDeepCrawlStrategy.can_process_url() and trace how discovered links reach filter_chain.apply(); reproduce the failure with https://name/xyz/ and a single-label hostname. Done means valid http/https URLs with a non-empty netloc are accepted and the configured DomainFilter is reached for links beyond depth 0.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

🐞 Bug 🩺 Needs Triage
crawl4ai version

0.9.2

Expected Behavior

Single-label/internal hostnames without a dot should be treated as valid
hostnames, at least when the URL has a valid scheme (http/https) and
non-empty netloc. The dot-check appears to be an overly strict heuristic
that isn't part of the documented API/config (no allow_single_label_hosts
or similar flag exists).

Current Behavior
  • Depth 0 (start URL) is fetched successfully.
  • Every discovered link at depth > 0 is logged as:
    Invalid URL: https://name/xyz/xyz, error: Invalid domain
    The configured DomainFilter is never even reached, since the exception
    is raised before filter_chain.apply() is called.
Is this reproducible?

Yes

Inputs Causing the Bug

Steps to Reproduce
1. Start URL: `https://name/xyz/` (a host with no dot in its name,
   reachable e.g. via internal DNS or /etc/hosts).
2. Configure `BFSDeepCrawlStrategy` with
   `filter_chain=FilterChain([DomainFilter(allowed_domains=["iltis3"])])`.
3. Run `crawler.arun()`.
Code snippets

OS

Linux

Python version

3.12.13

Browser

No response

Browser version

No response

Error logs & Screenshots (if applicable)

[INIT].... → Crawl4AI 0.9.2

[FETCH]... ↓ https://name/xyz | ✓ | ⏱: 0.91s

[SCRAPE].. ◆ https://name/xyz | ✓ | ⏱: 0.06s

[EXTRACT]. ■ https://name/xyz | ✓ | ⏱: 0.03s

[COMPLETE] ● https://name/xyz | ✓ | ⏱: 1.01s

Invalid URL: https://name/xyz/neu-im-intranet, error: Invalid domain

Invalid URL: name/xyz/suche, error: Invalid domain

Invalid URL: https://name/xyz

crawler.py

/allgemein/stellen, error: Invalid domain

Ngôn ngữ chính
Python
Star
84.5k
Fork
8.7k
Merge trung bình
3 ngày 9 giờ
Pull request đã merge (30 ngày)
17

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của unclecode/crawl4ai

Tất cả issue của unclecode/crawl4ai

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.