[Bug]: crawler.base_config boolean values are silently ignored (regression from #1505)
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 55/100
Hướng nghiên cứu
Start in api.py around CrawlerRunConfig.load() at line 675 and effective_config handling at lines 697-716; compare the normal and config-list paths, including the guard at line 715. Use the raw request dictionary to distinguish omitted fields from explicitly sent defaults. Done means base_config boolean and numeric defaults apply when omitted while explicit client values remain unchanged.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
crawl4ai version
0.9.2
Expected Behavior
config.yml sets a server-side default for every crawl:
# deploy/docker/config.yml:74-76
crawler:
base_config:
simulate_user: true
A POST /crawl request that does not send simulate_user should run with simulate_user=True, i.e. the server default applies.
A request that does send simulate_user should win over the server default — that is the intent of #1505.
Current Behavior
The server default never applies. simulate_user is False on every crawl.
api.py:715 reads "the client didn't send this field" as "the attribute is None or """:
current_value = getattr(crawler_config, key)
if current_value is None or current_value == "": # api.py:715
setattr(crawler_config, key, value)
CrawlerRunConfig.simulate_user defaults to False (async_configs.py:1650), and False is neither None nor "", so the guard never passes.
The same goes for every base_config key defaulting to a boolean or a number: magic, override_navigator, check_robots_txt, remove_overlay_elements, page_timeout.
The config-list path at api.py:707 (8995c1b, #1837) copies the guard.
So the stock image ships simulate_user: true (config.yml:74-76, utils.py:63)
but never injects the navigator_overrider script (browser_manager.py:1229-1235) or runs the mouse-move simulation (async_crawler_strategy.py:980-983).
a1950af (#1505) introduced this.
The setattr used to be unconditional and clobbered client-sent values, so reverting brings #1505 back.
After CrawlerRunConfig.load() (api.py:675) nothing tells "omitted" apart from "sent, equal to the default".
That information only exists in the raw request dict
Is this reproducible?
Yes
Inputs Causing the Bug
- URL(s): any, e.g. https://example.com
- Settings used: stock deploy/docker/config.yml, i.e. crawler.base_config.simulate_user: true
- Input data: {"urls": ["https://example.com"]} # no crawler_config key
Steps to Reproduce
1. Start the stock server image, config.yml untouched.
2. POST the body above to /crawl.
3. Read effective_config in handle_crawl_request (api.py:697-716).
simulate_user is False.
Code snippets
# The guard in isolation. No server or browser needed.
from crawl4ai import CrawlerRunConfig
cfg = CrawlerRunConfig() # client sent no crawler_config
value = getattr(cfg, "simulate_user") # False, the dataclass default
assert value is None or value == "" # api.py:715 -> fails, setattr skipped
OS
Linux (Docker image, python:3.12-slim-bookworm)
Python version
3.12
Browser
Chromium (Playwright, headless)
Browser version
No response
Error logs & Screenshots (if applicable)
No error. The server drops the value silently.
- Ngôn ngữ chính
- Python
- Star
- 84.5k
- Fork
- 8.7k
- Merge trung bình
- 3 ngày 9 giờ
- Pull request đã merge (30 ngày)
- 17
Chuẩn bị môi trường
- Có Dockerfile hoặc tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của unclecode/crawl4ai
-
[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh runĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
unclecode/crawl4ai#2309 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 84/100
unclecode/crawl4ai#2147 · 3 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
unclecode/crawl4ai#2123 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
🐞 Bug 🩺 Needs Triage
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 55/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của unclecode/crawl4ai
Issue tương tự
-
customer-reported
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
Azure/azure-cli#34150 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
community-request
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 95/100
NVIDIA-NeMo/Curator#2464 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
weblate-discover crashes with an unhandled FileNotFoundError when the directory does not existĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
WeblateOrg/translation-finder#1099 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
trezor/trezor-firmware#7997 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
Maintainer thường phản hồi trong vòng 1 ngày