[Bug]: MCP scrape tools lack wait_until / SPA support that REST API and CLI provide
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 52/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Tranquilla
- Stack tecnologico
- python
- Ambito
- api, documentation
Direzione di ricerca
Start at the MCP scrape tool schemas and compare their exposed arguments with the existing crawler_config options used by POST /crawl and crwl -c. Verify the MCP tools on a JavaScript-heavy page, then confirm they accept wait_until and delay_before_return_html and that the documentation explains default wait behavior across MCP, REST, and CLI.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
crawl4ai version
0.8.6
Expected Behavior
MCP scrape tools (crawl4ai_md, crawl4ai_html, etc.) should accept the same wait_until, delay_before_return_html, and cache_mode parameters as the REST API (POST /crawl) and CLI (crwl -c). This would allow MCP-based agents to reliably scrape JavaScript-heavy pages by waiting for content to fully load.
Example desired usage:
{
"url": "https://example.com/dynamic-content",
"wait_until": "networkidle",
"delay_before_return_html": 2
}
Environment
| Component | Version |
|---|---|
| Crawl4AI | 0.8.6 |
| OS | Debian GNU/Linux 12 (bookworm) |
| Python | 3.12.13 |
| Image | unclecode/crawl4ai:latest |
Suggested Fix
- Expose
crawler_configparameters on MCP tool schemas — mapwait_until,delay_before_return_html,cache_mode, etc. to the existing REST/CLI options - Document MCP defaults vs REST/CLI behavior
Acceptance Criteria
- MCP scrape tools accept
wait_untilparameter (load,domcontentloaded,networkidle,commit) - MCP scrape tools accept
delay_before_return_htmlparameter - Documentation clarifies default wait behavior for each interface (MCP vs REST vs CLI)
Current Behavior
MCP tools return immediately after initial DOM load, without waiting for dynamic content. No parameters are exposed to control wait behavior.
- REST API with
crawler_config.wait_until: "networkidle"→ ✅ Complete rendered content - CLI with
-c 'wait_until=networkidle'→ ✅ Complete rendered content - MCP tools → ❌ Incomplete content (dynamic elements missing)
Current workaround: Bypass MCP entirely and use POST /crawl directly.
Is this reproducible?
Yes
Inputs Causing the Bug
Any JavaScript-heavy / AJAX-driven page where content loads after initial page load.
Steps to Reproduce
1. Start Crawl4AI with MCP enabled at `/mcp/sse`
2. Configure any MCP client (e.g., OpenCode) with the Crawl4AI MCP server
3. Call MCP scrape tool:
crawl4ai_md(url="https://example.com/dynamic-content")
4. Observe: Content is incomplete (dynamic elements not rendered)
5. Compare with working REST call:
curl -s http://localhost:11235/crawl \
-H 'Content-Type: application/json' \
-d '{
"urls": ["https://example.com/dynamic-content"],
"crawler_config": {
"wait_until": "networkidle",
"cache_mode": "bypass"
}
}'
6. Observe: REST returns complete rendered content
Code snippets
OS
Debian GNU/Linux 12 (bookworm)
Python version
3.12.13
Browser
No response
Browser version
No response
Error logs & Screenshots (if applicable)
No response
- Lingua principale
- Python
- Stelle
- 84.5k
- Fork
- 8.7k
- Merge medio
- 3g 9h
- PR unite (30g)
- 17
Preparare l'ambiente
- Include un Dockerfile o un file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di unclecode/crawl4ai
-
[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh runAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
unclecode/crawl4ai#2309 · 2 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 84/100
unclecode/crawl4ai#2147 · 3 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
unclecode/crawl4ai#2123 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
🐞 Bug 🩺 Needs Triage
Difficoltà 4/5 3-5 giorni Idoneità per principianti 55/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di unclecode/crawl4ai
Issue simili
-
customer-reported
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
Azure/azure-cli#34150 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
community-request
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 95/100
NVIDIA-NeMo/Curator#2464 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
weblate-discover crashes with an unhandled FileNotFoundError when the directory does not existAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
WeblateOrg/translation-finder#1099 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
trezor/trezor-firmware#7997 ·
I maintainer di solito rispondono entro 2 giorni
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
I maintainer di solito rispondono entro 1 giorno