Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Config option to disable is_blocked() post-crawl content veto

Aperta
#2,058 2 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
55/100
Tipo di issue
Funzionalità
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
python
Ambito
backend

Direzione di ricerca

Start by tracing arun() post-processing in async_webcrawler.py and the is_blocked() implementation in antibot_detector.py, including the existing binary-download and fallback-fetch special cases. Then inspect CrawlerRunConfig and CrawlResult to determine the least disruptive opt-out or verdict exposure, and add regression coverage showing that fetched content remains available for caller-side quality checks.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

is_blocked() (antibot_detector.py) is applied unconditionally in arun()'s post-processing and can mark successfully-fetched content as failed. Your own code already special-cases it twice in async_webcrawler.py — the comments note it "would misread 0 bytes html as a block" for binary downloads, and that real pages containing anti-bot script markers "trigger false positives" after fallback fetches. We hit a third case: small legitimate pages (a 378-byte file:// document) fail the minimal_text structural tier. Downstream consumers with their own quality pipelines currently have no way to receive the fetched content plus the verdict instead of a hard failure. Would you accept a PR adding a CrawlerRunConfig flag (e.g. check_blocked: bool = True) — or alternatively surfacing the verdict on a successful CrawlResult field — so callers can opt into making that judgment themselves?

Lingua principale
Python
Stelle
84.5k
Fork
8.7k
Merge medio
3g 9h
PR unite (30g)
17

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di unclecode/crawl4ai

Tutte le issue di unclecode/crawl4ai

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.