Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[Bug]: MCP scrape tools lack wait_until / SPA support that REST API and CLI provide

オープン
#1,963 コメント 2 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
52/100
issue の種類
バグ
明瞭さ
おおむね明確
活発さ
静か
技術スタック
python
領域
api, documentation

調査の方向性

Start at the MCP scrape tool schemas and compare their exposed arguments with the existing crawler_config options used by POST /crawl and crwl -c. Verify the MCP tools on a JavaScript-heavy page, then confirm they accept wait_until and delay_before_return_html and that the documentation explains default wait behavior across MCP, REST, and CLI.

索引モデルが issue の本文から書いたものです。

説明

⚙️ In-progress 🐞 Bug 📌 Root caused
crawl4ai version

0.8.6

Expected Behavior

MCP scrape tools (crawl4ai_md, crawl4ai_html, etc.) should accept the same wait_until, delay_before_return_html, and cache_mode parameters as the REST API (POST /crawl) and CLI (crwl -c). This would allow MCP-based agents to reliably scrape JavaScript-heavy pages by waiting for content to fully load.

Example desired usage:

{
  "url": "https://example.com/dynamic-content",
  "wait_until": "networkidle",
  "delay_before_return_html": 2
}

Environment

Component Version
Crawl4AI 0.8.6
OS Debian GNU/Linux 12 (bookworm)
Python 3.12.13
Image unclecode/crawl4ai:latest

Suggested Fix

  1. Expose crawler_config parameters on MCP tool schemas — map wait_until, delay_before_return_html, cache_mode, etc. to the existing REST/CLI options
  2. Document MCP defaults vs REST/CLI behavior

Acceptance Criteria

  • MCP scrape tools accept wait_until parameter (load, domcontentloaded, networkidle, commit)
  • MCP scrape tools accept delay_before_return_html parameter
  • Documentation clarifies default wait behavior for each interface (MCP vs REST vs CLI)
Current Behavior

MCP tools return immediately after initial DOM load, without waiting for dynamic content. No parameters are exposed to control wait behavior.

  • REST API with crawler_config.wait_until: "networkidle" → ✅ Complete rendered content
  • CLI with -c 'wait_until=networkidle' → ✅ Complete rendered content
  • MCP tools → ❌ Incomplete content (dynamic elements missing)

Current workaround: Bypass MCP entirely and use POST /crawl directly.

Is this reproducible?

Yes

Inputs Causing the Bug
Any JavaScript-heavy / AJAX-driven page where content loads after initial page load.
Steps to Reproduce
1. Start Crawl4AI with MCP enabled at `/mcp/sse`
2. Configure any MCP client (e.g., OpenCode) with the Crawl4AI MCP server
3. Call MCP scrape tool:
   
   crawl4ai_md(url="https://example.com/dynamic-content")
   
4. Observe: Content is incomplete (dynamic elements not rendered)
5. Compare with working REST call:
   
   curl -s http://localhost:11235/crawl \
     -H 'Content-Type: application/json' \
     -d '{
       "urls": ["https://example.com/dynamic-content"],
       "crawler_config": {
         "wait_until": "networkidle",
         "cache_mode": "bypass"
       }
     }'
   
6. Observe: REST returns complete rendered content
Code snippets

OS

Debian GNU/Linux 12 (bookworm)

Python version

3.12.13

Browser

No response

Browser version

No response

Error logs & Screenshots (if applicable)

No response

主要言語
Python
スター
84.5k
フォーク
8.7k
平均マージ
3日 9時間
マージ済み PR(30日)
17

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

unclecode/crawl4ai のほかの issue

unclecode/crawl4ai の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。