[Bug]: RateLimiter provides ineffective protection against failures
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 35/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Ferma
- Stack tecnologico
- python
- Ambito
- backend, networking, performance
Direzione di ricerca
Start by locating the RateLimiter implementation and the deep-crawl dispatcher configuration. Reproduce crawling against gamesjobslive.niceboard.co and inspect request spacing, 429/503 responses, retries, and standard rate-limit headers. Done means rate-limited sites are processed successfully, deep crawls can configure the dispatcher, and the limiter responds to reported limits.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
crawl4ai version
6.0.0
Expected Behavior
A crawl should successful handle a site which actively manages client request rates.
Current Behavior
The current RateLimiter implementation uses a simple last request and current delay calculation, which could lead to uneven request distribution when multiple requests were made in quick succession.
The result of this is the more links discovered by as single request the more likely it is that we would trigger 429 and 503 response codes, when combined with max_retries this would cause the crawler to fail to successfully process all pages if the site implements rate limiting.
An example site: https://gamesjobslive.niceboard.co/
In addition to this it's currently not possible to configure the rate limiter for deep crawl as there is no way to set dispatcher.
Finally the rate limiter doesn't adapt to site which report their rate limits by the standard rate limiting headers, significantly increasing the number of retries and ultimately failures.
Is this reproducible?
Yes
- Lingua principale
- Python
- Stelle
- 84.5k
- Fork
- 8.7k
- Merge medio
- 3g 9h
- PR unite (30g)
- 17
Preparare l'ambiente
- Include un Dockerfile o un file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di unclecode/crawl4ai
-
[Bug]: Reusing BFSDeepCrawlStrategy leaks the previous crawl's max_pages budget into a fresh runAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
unclecode/crawl4ai#2309 · 2 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 84/100
unclecode/crawl4ai#2147 · 3 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
unclecode/crawl4ai#2123 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
🐞 Bug 🩺 Needs Triage
Difficoltà 4/5 3-5 giorni Idoneità per principianti 55/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di unclecode/crawl4ai
Issue simili
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 72/100
letsencrypt/cp-cps#353 ·
-
Marble Madness II is missingAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
PedestrianDynamics/pyFDS-Evac#394 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
DOI-USGS/pywatershed#421 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
python-pillow/Pillow#10087 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno