[Bug]: Failed CrawlResult has `links: {}` / `media: {}` instead of the usual keys
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 3/5
- Thời gian dự kiến
- 1-2 ngày
- Mức phù hợp với người mới
- 78/100
Hướng nghiên cứu
Start with CrawlResult defaults in crawl4ai/models.py and the result construction branches in crawl4ai/async_webcrawler.py, especially robots.txt refusal, proxy failure, exception, and anti-bot fallback paths. Reproduce the two-URL /crawl case with check_robots_txt=true, then verify that failed and fallback results expose the same internal/external link and image/video/audio media keys as successful results.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
crawl4ai version
0.9.4 (also on 0.9.2)
Expected Behavior
A result that failed (robots.txt refusal, fetch error, "All proxies failed") has the same links and media keys as a successful one:
"links": {"internal": [], "external": []},
"media": {"images": [], "videos": [], "audios": []}
so a client can read links["internal"] without special-casing failures.
Current Behavior
CrawlResult.links and CrawlResult.media default to {} (models.py#L136-L137). Successful crawls fill them from the scraper's typed Links / Media models, but the results AsyncWebCrawler.arun builds itself never set them:
- robots.txt refusal (async_webcrawler.py#L385-L397)
- "All proxies failed" (#L640-L646)
- the exception path (#L710-L712)
- the anti-bot raw-HTML fallback, which is a successful result (#L588-L598)
So in a mixed batch /crawl returns "links": {}, "media": {} for exactly the entries a client has to handle differently, and in the SDK result.links["internal"] raises KeyError. Since #2134 keeps failed results in /crawl responses, every batch with a refused or failed URL has this shape. In one client that read links.internal.length, the resulting TypeError was taken for a failed request and the healthy pages of the batch were re-crawled several times over.
cc @nightcityblade (#2134) @ntohidi
Is this reproducible?
Yes
Inputs Causing the Bug
- Two URLs on one site, one of them disallowed by its robots.txt
- crawler_config params: {"check_robots_txt": true}
Steps to Reproduce
1. Run the Docker image (0.9.4, or built from develop)
2. POST /crawl with both URLs and check_robots_txt=true
3. Compare `links` / `media` of the two results
Code snippets
from crawl4ai.models import CrawlResult
# What AsyncWebCrawler.arun returns for a URL robots.txt disallows
r = CrawlResult(url="https://example.com/private", html="", success=False,
status_code=403, error_message="Access denied by robots.txt")
print(r.links, r.media) # {} {}
r.links["internal"] # KeyError: 'internal'
OS
Linux (Docker image built from develop @ 1f68e5b)
Python version
3.12 (the image's)
Browser
Chromium (Playwright, bundled)
Browser version
No response
Error logs & Screenshots (if applicable)
POST /crawl, check_robots_txt=true, two URLs of a local test site (CRAWL4AI_ALLOW_INTERNAL_URLS=true), /private disallowed by its robots.txt:
[shape] HTTP 200
/index.html success=True status=200 links=['external', 'internal'] media=['audios', 'images', 'videos'] error=''
/private/index.html success=False status=403 links={} media={} error='Access denied by robots.txt'
- Ngôn ngữ chính
- Python
- Star
- 84.5k
- Fork
- 8.7k
- Merge trung bình
- 3 ngày 9 giờ
- Pull request đã merge (30 ngày)
- 17
Chuẩn bị môi trường
- Có Dockerfile hoặc tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của unclecode/crawl4ai
-
[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh runĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
unclecode/crawl4ai#2309 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 84/100
unclecode/crawl4ai#2147 · 3 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
unclecode/crawl4ai#2123 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
🐞 Bug 🩺 Needs Triage
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 55/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của unclecode/crawl4ai
Issue tương tự
-
bug server
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
sportsdataverse/sportsdataverse-py#641 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
googleapis/google-cloud-python#18532 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
Maintainer thường phản hồi trong vòng 1 ngày