Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

[Bug]: Failed CrawlResult has `links: {}` / `media: {}` instead of the usual keys

未关闭
#2,286 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
3/5
预计耗时
1-2 天
新手友好度
78/100
Issue 类型
缺陷
描述清晰度
描述清楚
活跃度
活跃
技术栈
python
领域
api, backend

调研方向

Start with CrawlResult defaults in crawl4ai/models.py and the result construction branches in crawl4ai/async_webcrawler.py, especially robots.txt refusal, proxy failure, exception, and anti-bot fallback paths. Reproduce the two-URL /crawl case with check_robots_txt=true, then verify that failed and fallback results expose the same internal/external link and image/video/audio media keys as successful results.

由索引模型根据 Issue 内容生成。

描述

🐞 Bug 🩺 Needs Triage
crawl4ai version

0.9.4 (also on 0.9.2)

Expected Behavior

A result that failed (robots.txt refusal, fetch error, "All proxies failed") has the same links and media keys as a successful one:

"links": {"internal": [], "external": []},
"media": {"images": [], "videos": [], "audios": []}

so a client can read links["internal"] without special-casing failures.

Current Behavior

CrawlResult.links and CrawlResult.media default to {} (models.py#L136-L137). Successful crawls fill them from the scraper's typed Links / Media models, but the results AsyncWebCrawler.arun builds itself never set them:

So in a mixed batch /crawl returns "links": {}, "media": {} for exactly the entries a client has to handle differently, and in the SDK result.links["internal"] raises KeyError. Since #2134 keeps failed results in /crawl responses, every batch with a refused or failed URL has this shape. In one client that read links.internal.length, the resulting TypeError was taken for a failed request and the healthy pages of the batch were re-crawled several times over.

cc @nightcityblade (#2134) @ntohidi

Is this reproducible?

Yes

Inputs Causing the Bug
- Two URLs on one site, one of them disallowed by its robots.txt
- crawler_config params: {"check_robots_txt": true}
Steps to Reproduce
1. Run the Docker image (0.9.4, or built from develop)
2. POST /crawl with both URLs and check_robots_txt=true
3. Compare `links` / `media` of the two results
Code snippets
from crawl4ai.models import CrawlResult

# What AsyncWebCrawler.arun returns for a URL robots.txt disallows
r = CrawlResult(url="https://example.com/private", html="", success=False,
                status_code=403, error_message="Access denied by robots.txt")
print(r.links, r.media)   # {} {}
r.links["internal"]       # KeyError: 'internal'
OS

Linux (Docker image built from develop @ 1f68e5b)

Python version

3.12 (the image's)

Browser

Chromium (Playwright, bundled)

Browser version

No response

Error logs & Screenshots (if applicable)

POST /crawl, check_robots_txt=true, two URLs of a local test site (CRAWL4AI_ALLOW_INTERNAL_URLS=true), /private disallowed by its robots.txt:
[shape] HTTP 200
/index.html success=True status=200 links=['external', 'internal'] media=['audios', 'images', 'videos'] error=''
/private/index.html success=False status=403 links={} media={} error='Access denied by robots.txt'

主要语言
Python
星标
84.5k
派生
8.7k
平均合并
3 天 20 小时
30 天内合并 PR
15

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

unclecode/crawl4ai 的其他 Issue

查看 unclecode/crawl4ai 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。