Hacktoberfest 2026: die Issues, die Maintainer für den Oktober markiert haben – offen und einsteigerfreundlich. Hacktoberfest-Issues durchsuchen

Relative links in a page whose URL contains `//` are resolved with the slashes collapsed

Offen
#346 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Dieses Issue hat noch niemand übernommen.

Bewertung

Schwierigkeit
3/5
Geschätzter Aufwand
1-2 Tage
Anfängerfreundlichkeit
68/100
Issue-Typ
Bug
Klarheit
Größtenteils klar
Aktivitätsstatus
Aktiv
Tech-Stack
python
Bereich
backend

Rechercherichtung

Start in ArticleUrlRewriter.__call__, where the issue shows nested urllib.parse.urljoin calls for article_url, base_href, and item_url; read the resolver tests from #344 for the existing approach. Check both relative-link and base_href resolution with URLs containing consecutive slashes, and confirm the resulting ZIM path preserves the empty segments.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Beschreibung

Follow-up to #344 / #340.

When a page's own URL contains consecutive slashes, relative links inside it are resolved with urllib.parse.urljoin, which drops the empty segments. So the // gets collapsed before normalize() even sees the URL.

Example (from @Sinkleberg's check in #344): rewriting other.html from https://example.com/x//y/page.html gives the ZIM path example.com/x/y/other.html, but a browser (WHATWG URL resolution) resolves it to https://example.com/x//y/other.html. So the link points to an entry that doesn't exist.

This is in ArticleUrlRewriter.__call__:

item_absolute_url = urljoin(
    urljoin(self.article_url.value, base_href), item_url
)

The same applies to base_href resolution. A fix probably needs an RFC 3986 / WHATWG-style join that keeps empty segments instead of urljoin. #344 already has a small resolver like that in its tests.

Vorherrschende Sprache
Python
Sterne
31
Forks
31
Ø Merge
2 T. 5 Std.
Gemergte PRs (30 T.)
3

Entwicklungsumgebung

Dieses Projekt bietet weder Dev-Container noch Dockerfile noch Beitragsleitfaden – die Einrichtung liegt bei Ihnen. Beginnen Sie mit der README; die allgemeinen Schritte stehen in unserem Leitfaden für den ersten Beitrag.

Erste Schritte

  1. Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
  3. Forken Sie das Repository und arbeiten Sie in einem Branch.
  4. Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.

Mehr aus openzim/python-scraperlib

Alle Issues in openzim/python-scraperlib

Ähnliche Issues

Weitere Issues zu Python

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.