[Bug]: RateLimiter provides ineffective protection against failures
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 35/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Đình trệ
- Công nghệ
- python
- Lĩnh vực
- backend, networking, performance
Hướng nghiên cứu
Start by locating the RateLimiter implementation and the deep-crawl dispatcher configuration. Reproduce crawling against gamesjobslive.niceboard.co and inspect request spacing, 429/503 responses, retries, and standard rate-limit headers. Done means rate-limited sites are processed successfully, deep crawls can configure the dispatcher, and the limiter responds to reported limits.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
crawl4ai version
6.0.0
Expected Behavior
A crawl should successful handle a site which actively manages client request rates.
Current Behavior
The current RateLimiter implementation uses a simple last request and current delay calculation, which could lead to uneven request distribution when multiple requests were made in quick succession.
The result of this is the more links discovered by as single request the more likely it is that we would trigger 429 and 503 response codes, when combined with max_retries this would cause the crawler to fail to successfully process all pages if the site implements rate limiting.
An example site: https://gamesjobslive.niceboard.co/
In addition to this it's currently not possible to configure the rate limiter for deep crawl as there is no way to set dispatcher.
Finally the rate limiter doesn't adapt to site which report their rate limits by the standard rate limiting headers, significantly increasing the number of retries and ultimately failures.
Is this reproducible?
Yes
- Ngôn ngữ chính
- Python
- Star
- 84.5k
- Fork
- 8.7k
- Merge trung bình
- 3 ngày 9 giờ
- Pull request đã merge (30 ngày)
- 17
Chuẩn bị môi trường
- Có Dockerfile hoặc tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của unclecode/crawl4ai
-
[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh runĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 84/100
unclecode/crawl4ai#2147 · 3 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
unclecode/crawl4ai#2123 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
🐞 Bug 🩺 Needs Triage
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 55/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của unclecode/crawl4ai
Issue tương tự
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 90/100
Maintainer thường phản hồi trong vòng 1 ngày
-
https://search.utilibre.orgĐang mởinstance instance add
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
searxng/searx-instances#941 · 1 bình luận ·
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
FluidNumerics/fluid-walk-blocker#89 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
Maintainer thường phản hồi trong vòng 1 ngày