Relative links in a page whose URL contains `//` are resolved with the slashes collapsed
還沒有人認領這個 Issue。
評估
研究方向
Start in ArticleUrlRewriter.__call__, where the issue shows nested urllib.parse.urljoin calls for article_url, base_href, and item_url; read the resolver tests from #344 for the existing approach. Check both relative-link and base_href resolution with URLs containing consecutive slashes, and confirm the resulting ZIM path preserves the empty segments.
由索引模型根據 Issue 內容生成。
描述
Follow-up to #344 / #340.
When a page's own URL contains consecutive slashes, relative links inside it are resolved with urllib.parse.urljoin, which drops the empty segments. So the // gets collapsed before normalize() even sees the URL.
Example (from @Sinkleberg's check in #344): rewriting other.html from https://example.com/x//y/page.html gives the ZIM path example.com/x/y/other.html, but a browser (WHATWG URL resolution) resolves it to https://example.com/x//y/other.html. So the link points to an entry that doesn't exist.
This is in ArticleUrlRewriter.__call__:
item_absolute_url = urljoin(
urljoin(self.article_url.value, base_href), item_url
)
The same applies to base_href resolution. A fix probably needs an RFC 3986 / WHATWG-style join that keeps empty segments instead of urljoin. #344 already has a small resolver like that in its tests.
- 主要語言
- Python
- 星號
- 31
- 分支
- 31
- 平均合併
- 2 天 5 小時
- 30 天內合併 PR
- 3
環境準備
這個專案沒有提供開發容器、Dockerfile 或貢獻指南,環境需要你自己搭建:先看它的 README,通用步驟見我們的新手貢獻指南。
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
openzim/python-scraperlib 的其他 Issue
-
HTML rewriting: also rewrite `poster` attribute可能已有人在做 關聯的 PR 仍在進行中或已合併。 未關閉
難度 2/5 1-3 小時 新手友好度 72/100
openzim/python-scraperlib#339 · 1 則留言 ·
-
難度 2/5 1-3 小時 新手友好度 68/100
openzim/python-scraperlib#292 ·
-
URL normalisation: do not rewrite consecutive slashes `//` as a single slash `/`可能已有人在做 @anshuman83-40 於 7 天前認領。 未關閉
難度 3/5 1-2 天 新手友好度 45/100
openzim/python-scraperlib#340 ·
-
Add fuzzy rule to rewrite URLs of lesbases.anct.gouv.fr可能重新可做 @benoit74 於 51 天前認領,目前沒有進行中的 PR。 未關閉
openzim/python-scraperlib#334 · 已指派 1 人 ·
-
Add utility to index ePub documents content可能已有人在做 @Sriram-PR 於 5 天前認領。 未關閉
難度 3/5 1-2 天 新手友好度 55/100
openzim/python-scraperlib#333 · 1 則留言 ·
查看 openzim/python-scraperlib 的全部 Issue
相似的 Issue
-
namespace operations
難度 1/5 1 小時以內 新手友好度 72/100
EclipseFdn/open-vsx.org#14043 ·
維護者通常 1 天內回覆
-
netbox status: needs triage type: bug
難度 2/5 1-3 小時 新手友好度 76/100
netbox-community/netbox#23376 ·
維護者通常 1 天內回覆
-
feedback simulation workshop
難度 2/5 1-3 小時 新手友好度 73/100
githubnext/gh-aw-workshop#4455 ·
維護者通常 1 天內回覆
-
Triage 🩺
難度 2/5 1-3 小時 新手友好度 76/100
維護者通常 1 天內回覆
-
[BUG] Container scenario crashes without expected_recovery_time, kube DNS example uses retry_wait未關閉needs-triage
難度 2/5 1-3 小時 新手友好度 77/100
krkn-chaos/krkn#1627 · 1 則留言 ·
維護者通常 1 天內回覆