Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

[Bug]: One seed that fails the destination check rejects the whole `/crawl` batch

Abierto
#2,288 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
3/5
Tiempo estimado
1-2 días
Aptitud para principiantes
72/100
Tipo de issue
Error
Claridad
Bastante claro
Estado de actividad
Activo
Stack tecnológico
python
Área
api, backend, security

Línea de trabajo

Start at _normalize_and_validate_seeds in deploy/docker/api.py around lines 650-657 and trace how /crawl collects results. Then inspect resolve_and_pin in deploy/docker/egress_broker.py around lines 92-96, reproducing the two-seed request from the issue. Done means one refused or unresolvable seed becomes an opaque failed result while other seeds still crawl successfully.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

🐞 Bug 🩺 Needs Triage
crawl4ai version

0.9.4

Expected Behavior

The other seeds are crawled, and the refused one comes back as a failed result carrying the same opaque message, the way a robots.txt refusal already does inside a batch.

Current Behavior

_normalize_and_validate_seeds checks every seed before crawling and raises on the first one refused (api.py#L650-L657), so the whole request fails with 400 {"detail": "URL blocked (SSRF protection): URL blocked"} and no results. resolve_and_pin also raises EgressBlocked when getaddrinfo fails (egress_broker.py#L92-L96), so a host that simply doesn't resolve (a dead domain from a stale sitemap) takes the batch down the same way. The caller can't tell which seed it was, so its only recovery is to split the batch and retry.

Keeping the message opaque, with no resolution oracle, is right, and a per-URL failure keeps it: it says no more than today's 400, because a caller can already find the refused seed by sending the URLs one at a time.

If per-URL failures are the direction you want, I can send a PR. cc @ntohidi

Is this reproducible?

Yes

Inputs Causing the Bug
- urls: ["https://example.com/", "https://no-such-host-12345.example/"]
Steps to Reproduce
1. Run the Docker image (0.9.4, or built from develop), without CRAWL4AI_ALLOW_INTERNAL_URLS
2. POST /crawl with the two URLs above
3. 400, no results; the first URL alone crawls fine
Code snippets
import httpx

r = httpx.post("http://localhost:11235/crawl", headers={"Authorization": f"Bearer {TOKEN}"},
               json={"urls": ["https://example.com/", "https://no-such-host-12345.example/"]}, timeout=60)
print(r.status_code, r.text)   # 400 {"detail":"URL blocked (SSRF protection): URL blocked"}
OS

Linux (Docker image built from develop @ 1f68e5b)

Python version

3.12 (the image's)

Browser

Chromium (Playwright, bundled)

Browser version

No response

Error logs & Screenshots (if applicable)

[seedbatch] one resolvable + one unresolvable seed -> HTTP 400: {"detail":"URL blocked (SSRF protection): URL blocked"}
[seedbatch] the resolvable seed alone -> HTTP 200, success=True

Lenguaje dominante
Python
Estrellas
84.5k
Forks
8.7k
Merge medio
3 d 20 h
PR fusionados (30 d)
15

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de unclecode/crawl4ai

Todos los issues de unclecode/crawl4ai

Issues similares

Más issues de Python

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.