Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

[Bug]: Failed CrawlResult has `links: {}` / `media: {}` instead of the usual keys

Đang mở
#2,286 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
3/5
Thời gian dự kiến
1-2 ngày
Mức phù hợp với người mới
78/100
Loại issue
Lỗi
Độ rõ ràng
Đặc tả rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
python
Lĩnh vực
api, backend

Hướng nghiên cứu

Start with CrawlResult defaults in crawl4ai/models.py and the result construction branches in crawl4ai/async_webcrawler.py, especially robots.txt refusal, proxy failure, exception, and anti-bot fallback paths. Reproduce the two-URL /crawl case with check_robots_txt=true, then verify that failed and fallback results expose the same internal/external link and image/video/audio media keys as successful results.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

🐞 Bug 🩺 Needs Triage
crawl4ai version

0.9.4 (also on 0.9.2)

Expected Behavior

A result that failed (robots.txt refusal, fetch error, "All proxies failed") has the same links and media keys as a successful one:

"links": {"internal": [], "external": []},
"media": {"images": [], "videos": [], "audios": []}

so a client can read links["internal"] without special-casing failures.

Current Behavior

CrawlResult.links and CrawlResult.media default to {} (models.py#L136-L137). Successful crawls fill them from the scraper's typed Links / Media models, but the results AsyncWebCrawler.arun builds itself never set them:

So in a mixed batch /crawl returns "links": {}, "media": {} for exactly the entries a client has to handle differently, and in the SDK result.links["internal"] raises KeyError. Since #2134 keeps failed results in /crawl responses, every batch with a refused or failed URL has this shape. In one client that read links.internal.length, the resulting TypeError was taken for a failed request and the healthy pages of the batch were re-crawled several times over.

cc @nightcityblade (#2134) @ntohidi

Is this reproducible?

Yes

Inputs Causing the Bug
- Two URLs on one site, one of them disallowed by its robots.txt
- crawler_config params: {"check_robots_txt": true}
Steps to Reproduce
1. Run the Docker image (0.9.4, or built from develop)
2. POST /crawl with both URLs and check_robots_txt=true
3. Compare `links` / `media` of the two results
Code snippets
from crawl4ai.models import CrawlResult

# What AsyncWebCrawler.arun returns for a URL robots.txt disallows
r = CrawlResult(url="https://example.com/private", html="", success=False,
                status_code=403, error_message="Access denied by robots.txt")
print(r.links, r.media)   # {} {}
r.links["internal"]       # KeyError: 'internal'
OS

Linux (Docker image built from develop @ 1f68e5b)

Python version

3.12 (the image's)

Browser

Chromium (Playwright, bundled)

Browser version

No response

Error logs & Screenshots (if applicable)

POST /crawl, check_robots_txt=true, two URLs of a local test site (CRAWL4AI_ALLOW_INTERNAL_URLS=true), /private disallowed by its robots.txt:
[shape] HTTP 200
/index.html success=True status=200 links=['external', 'internal'] media=['audios', 'images', 'videos'] error=''
/private/index.html success=False status=403 links={} media={} error='Access denied by robots.txt'

Ngôn ngữ chính
Python
Star
84.5k
Fork
8.7k
Merge trung bình
3 ngày 9 giờ
Pull request đã merge (30 ngày)
17

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của unclecode/crawl4ai

Tất cả issue của unclecode/crawl4ai

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.