Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Config option to disable is_blocked() post-crawl content veto

Abierto
#2,058 2 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

@nightcityblade ya está trabajando en esto.

Desde el 19/7/2026.

Evaluación

Dificultad
4/5
Tiempo estimado
3-5 días
Aptitud para principiantes
55/100
Tipo de issue
Nueva funcionalidad
Claridad
Bastante claro
Estado de actividad
Activo
Stack tecnológico
python
Área
backend

Línea de trabajo

Start by tracing arun() post-processing in async_webcrawler.py and the is_blocked() implementation in antibot_detector.py, including the existing binary-download and fallback-fetch special cases. Then inspect CrawlerRunConfig and CrawlResult to determine the least disruptive opt-out or verdict exposure, and add regression coverage showing that fetched content remains available for caller-side quality checks.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

is_blocked() (antibot_detector.py) is applied unconditionally in arun()'s post-processing and can mark successfully-fetched content as failed. Your own code already special-cases it twice in async_webcrawler.py — the comments note it "would misread 0 bytes html as a block" for binary downloads, and that real pages containing anti-bot script markers "trigger false positives" after fallback fetches. We hit a third case: small legitimate pages (a 378-byte file:// document) fail the minimal_text structural tier. Downstream consumers with their own quality pipelines currently have no way to receive the fetched content plus the verdict instead of a hard failure. Would you accept a PR adding a CrawlerRunConfig flag (e.g. check_blocked: bool = True) — or alternatively surfacing the verdict on a successful CrawlResult field — so callers can opt into making that judgment themselves?

Lenguaje dominante
Python
Estrellas
84.8k
Forks
8.8k
Merge medio
13 h 57 min
PR fusionados (30 d)
13

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de unclecode/crawl4ai

Todos los issues de unclecode/crawl4ai

Issues similares

Más issues de Python

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.