Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

[Bug]: MCP scrape tools lack wait_until / SPA support that REST API and CLI provide

未关闭
#1,963 2 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
52/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
冷清
技术栈
python
领域
api, documentation

调研方向

Start at the MCP scrape tool schemas and compare their exposed arguments with the existing crawler_config options used by POST /crawl and crwl -c. Verify the MCP tools on a JavaScript-heavy page, then confirm they accept wait_until and delay_before_return_html and that the documentation explains default wait behavior across MCP, REST, and CLI.

由索引模型根据 Issue 内容生成。

描述

⚙️ In-progress 🐞 Bug 📌 Root caused
crawl4ai version

0.8.6

Expected Behavior

MCP scrape tools (crawl4ai_md, crawl4ai_html, etc.) should accept the same wait_until, delay_before_return_html, and cache_mode parameters as the REST API (POST /crawl) and CLI (crwl -c). This would allow MCP-based agents to reliably scrape JavaScript-heavy pages by waiting for content to fully load.

Example desired usage:

{
  "url": "https://example.com/dynamic-content",
  "wait_until": "networkidle",
  "delay_before_return_html": 2
}

Environment

Component Version
Crawl4AI 0.8.6
OS Debian GNU/Linux 12 (bookworm)
Python 3.12.13
Image unclecode/crawl4ai:latest

Suggested Fix

  1. Expose crawler_config parameters on MCP tool schemas — map wait_until, delay_before_return_html, cache_mode, etc. to the existing REST/CLI options
  2. Document MCP defaults vs REST/CLI behavior

Acceptance Criteria

  • MCP scrape tools accept wait_until parameter (load, domcontentloaded, networkidle, commit)
  • MCP scrape tools accept delay_before_return_html parameter
  • Documentation clarifies default wait behavior for each interface (MCP vs REST vs CLI)
Current Behavior

MCP tools return immediately after initial DOM load, without waiting for dynamic content. No parameters are exposed to control wait behavior.

  • REST API with crawler_config.wait_until: "networkidle" → ✅ Complete rendered content
  • CLI with -c 'wait_until=networkidle' → ✅ Complete rendered content
  • MCP tools → ❌ Incomplete content (dynamic elements missing)

Current workaround: Bypass MCP entirely and use POST /crawl directly.

Is this reproducible?

Yes

Inputs Causing the Bug
Any JavaScript-heavy / AJAX-driven page where content loads after initial page load.
Steps to Reproduce
1. Start Crawl4AI with MCP enabled at `/mcp/sse`
2. Configure any MCP client (e.g., OpenCode) with the Crawl4AI MCP server
3. Call MCP scrape tool:
   
   crawl4ai_md(url="https://example.com/dynamic-content")
   
4. Observe: Content is incomplete (dynamic elements not rendered)
5. Compare with working REST call:
   
   curl -s http://localhost:11235/crawl \
     -H 'Content-Type: application/json' \
     -d '{
       "urls": ["https://example.com/dynamic-content"],
       "crawler_config": {
         "wait_until": "networkidle",
         "cache_mode": "bypass"
       }
     }'
   
6. Observe: REST returns complete rendered content
Code snippets

OS

Debian GNU/Linux 12 (bookworm)

Python version

3.12.13

Browser

No response

Browser version

No response

Error logs & Screenshots (if applicable)

No response

主要语言
Python
星标
84.5k
派生
8.7k
平均合并
3 天 9 小时
30 天内合并 PR
17

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

unclecode/crawl4ai 的其他 Issue

查看 unclecode/crawl4ai 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。