[Bug]: `crawler_configs` is silently ignored for a single URL and on `/crawl/stream`
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 3/5
- Thời gian dự kiến
- 1-2 ngày
- Mức phù hợp với người mới
- 76/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Đặc tả rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- python
- Lĩnh vực
- api, backend, testing-qa
Hướng nghiên cứu
Start with handle_crawl_request in deploy/docker/api.py and compare its single-URL and multi-URL paths, then trace stream_process in deploy/docker/server.py into handle_stream_crawl_request. Run tests/test_issue_1837_config_list.py::test_single_url_ignores_crawler_configs and reproduce the documented /crawl and /crawl/stream cases. Done means crawler_configs are honored or unsupported cases return 400 instead of silently returning 200.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
crawl4ai version
0.9.4
Expected Behavior
crawler_configs applies to every request that carries it: a /crawl with one URL, and /crawl/stream (or /crawl with stream: true). If some case can't support it, a 400 rather than a 200 that ran with a different config.
Current Behavior
Two gaps in the config-list support from #1852:
handle_crawl_requestuses the list only whenlen(urls) > 1(api.py#L726); with one URL it callsarun()withcrawler_config(#L749).stream_processdoesn't passcrawler_configsdown, andhandle_stream_crawl_requesthas no parameter for it (server.py#L1030-L1036, api.py#L889-L895).
Both answer 200, so the caller can't tell its per-URL settings were dropped. The single-URL case is pinned by tests/test_issue_1837_config_list.py::test_single_url_ignores_crawler_configs ("arun only takes one config"), but arun_many handles a one-URL list fine.
cc @hafezparast @ntohidi
Is this reproducible?
Yes
Inputs Causing the Bug
- /crawl, one URL, crawler_configs: [{url_matcher: "*", css_selector: ".only-a"}]
- /crawl/stream, two URLs, the same list
Steps to Reproduce
1. Page with two sections, .only-a and .only-b
2. POST /crawl with that one URL and the crawler_configs above
3. The markdown still contains the .only-b section
4. Same list on /crawl/stream: same result
Code snippets
import httpx
# URL: a page with
# <div class="only-a"><p>SECTION-A is the part a per-URL css_selector keeps.</p></div>
# <div class="only-b"><p>SECTION-B is the part it drops.</p></div>
per_url = [{"type": "CrawlerRunConfig",
"params": {"url_matcher": "*", "css_selector": ".only-a", "cache_mode": "bypass"}}]
r = httpx.post("http://localhost:11235/crawl", headers={"Authorization": f"Bearer {TOKEN}"},
json={"urls": [URL], "crawler_configs": per_url}, timeout=60)
md = r.json()["results"][0]["markdown"]["raw_markdown"]
print("SECTION-B" in md) # True: the per-URL css_selector was not applied
OS
Linux (Docker image built from develop @ 1f68e5b)
Python version
3.12 (the image's)
Browser
Chromium (Playwright, bundled)
Browser version
No response
Error logs & Screenshots (if applicable)
Local test page with .only-a / .only-b sections (CRAWL4AI_ALLOW_INTERNAL_URLS=true); A / B = that section is in the markdown:
[configs] /crawl, one URL, per-URL css_selector=.only-a -> HTTP 200: A=True B=True
[configs] /crawl/stream, two URLs, same list -> HTTP 200: A=True B=True; A=True B=True
- Ngôn ngữ chính
- Python
- Star
- 84.5k
- Fork
- 8.7k
- Merge trung bình
- 3 ngày 9 giờ
- Pull request đã merge (30 ngày)
- 17
Chuẩn bị môi trường
- Có Dockerfile hoặc tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của unclecode/crawl4ai
-
[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh runĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
unclecode/crawl4ai#2309 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 84/100
unclecode/crawl4ai#2147 · 3 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
unclecode/crawl4ai#2123 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
🐞 Bug 🩺 Needs Triage
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 55/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của unclecode/crawl4ai
Issue tương tự
-
customer-reported
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
Azure/azure-cli#34150 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
community-request
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 95/100
NVIDIA-NeMo/Curator#2464 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
weblate-discover crashes with an unhandled FileNotFoundError when the directory does not existĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
WeblateOrg/translation-finder#1099 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
trezor/trezor-firmware#7997 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
Maintainer thường phản hồi trong vòng 1 ngày