Hacktoberfest 2026:維護者為十月標記出來的 issue,仍然開放、適合新手。 瀏覽 Hacktoberfest issue

Relative links in a page whose URL contains `//` are resolved with the slashes collapsed

未關閉
#346 0 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

還沒有人認領這個 Issue。

評估

難度
3/5
預估耗時
1-2 天
新手友好度
68/100
Issue 類型
缺陷
描述清晰度
基本清楚
活躍度
活躍
技術堆疊
python
領域
backend

研究方向

Start in ArticleUrlRewriter.__call__, where the issue shows nested urllib.parse.urljoin calls for article_url, base_href, and item_url; read the resolver tests from #344 for the existing approach. Check both relative-link and base_href resolution with URLs containing consecutive slashes, and confirm the resulting ZIM path preserves the empty segments.

由索引模型根據 Issue 內容生成。

描述

Follow-up to #344 / #340.

When a page's own URL contains consecutive slashes, relative links inside it are resolved with urllib.parse.urljoin, which drops the empty segments. So the // gets collapsed before normalize() even sees the URL.

Example (from @Sinkleberg's check in #344): rewriting other.html from https://example.com/x//y/page.html gives the ZIM path example.com/x/y/other.html, but a browser (WHATWG URL resolution) resolves it to https://example.com/x//y/other.html. So the link points to an entry that doesn't exist.

This is in ArticleUrlRewriter.__call__:

item_absolute_url = urljoin(
    urljoin(self.article_url.value, base_href), item_url
)

The same applies to base_href resolution. A fix probably needs an RFC 3986 / WHATWG-style join that keeps empty segments instead of urljoin. #344 already has a small resolver like that in its tests.

主要語言
Python
星號
31
分支
31
平均合併
2 天 5 小時
30 天內合併 PR
3

環境準備

這個專案沒有提供開發容器、Dockerfile 或貢獻指南,環境需要你自己搭建:先看它的 README,通用步驟見我們的新手貢獻指南。

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

openzim/python-scraperlib 的其他 Issue

查看 openzim/python-scraperlib 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。