Relative links in a page whose URL contains `//` are resolved with the slashes collapsed
还没有人认领这个 Issue。
评估
调研方向
Start in ArticleUrlRewriter.__call__, where the issue shows nested urllib.parse.urljoin calls for article_url, base_href, and item_url; read the resolver tests from #344 for the existing approach. Check both relative-link and base_href resolution with URLs containing consecutive slashes, and confirm the resulting ZIM path preserves the empty segments.
由索引模型根据 Issue 内容生成。
描述
Follow-up to #344 / #340.
When a page's own URL contains consecutive slashes, relative links inside it are resolved with urllib.parse.urljoin, which drops the empty segments. So the // gets collapsed before normalize() even sees the URL.
Example (from @Sinkleberg's check in #344): rewriting other.html from https://example.com/x//y/page.html gives the ZIM path example.com/x/y/other.html, but a browser (WHATWG URL resolution) resolves it to https://example.com/x//y/other.html. So the link points to an entry that doesn't exist.
This is in ArticleUrlRewriter.__call__:
item_absolute_url = urljoin(
urljoin(self.article_url.value, base_href), item_url
)
The same applies to base_href resolution. A fix probably needs an RFC 3986 / WHATWG-style join that keeps empty segments instead of urljoin. #344 already has a small resolver like that in its tests.
- 主要语言
- Python
- 星标
- 31
- 派生
- 31
- 平均合并
- 2 天 5 小时
- 30 天内合并 PR
- 3
环境准备
这个项目没有提供开发容器、Dockerfile 或贡献指南,环境需要你自己搭建:先看它的 README,通用步骤见我们的新手贡献指南。
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
openzim/python-scraperlib 的其他 Issue
-
HTML rewriting: also rewrite `poster` attribute可能已有人在做 关联的 PR 仍在进行中或已合并。 未关闭
难度 2/5 1-3 小时 新手友好度 72/100
openzim/python-scraperlib#339 · 1 条评论 ·
-
难度 2/5 1-3 小时 新手友好度 68/100
openzim/python-scraperlib#292 ·
-
URL normalisation: do not rewrite consecutive slashes `//` as a single slash `/`可能已有人在做 @anshuman83-40 于 7 天前认领。 未关闭
难度 3/5 1-2 天 新手友好度 45/100
openzim/python-scraperlib#340 ·
-
Add fuzzy rule to rewrite URLs of lesbases.anct.gouv.fr可能重新可做 @benoit74 于 51 天前认领,目前没有进行中的 PR。 未关闭
openzim/python-scraperlib#334 · 已指派 1 人 ·
-
Add utility to index ePub documents content可能已有人在做 @Sriram-PR 于 5 天前认领。 未关闭
难度 3/5 1-2 天 新手友好度 55/100
openzim/python-scraperlib#333 · 1 条评论 ·
查看 openzim/python-scraperlib 的全部 Issue
相似的 Issue
-
难度 1/5 1 小时以内 新手友好度 60/100
521xueweihan/HelloGitHub#3924 ·
-
难度 2/5 1-3 小时 新手友好度 67/100
wilbowes/EchoMuse#869 · 1 条评论 ·
维护者通常 1 天内回复
-
难度 1/5 1 小时以内 新手友好度 85/100
-
namespace operations
难度 1/5 1 小时以内 新手友好度 72/100
EclipseFdn/open-vsx.org#14043 ·
维护者通常 1 天内回复
-
test: TestServeUntilStale races the server's close against the client's sendall (BrokenPipeError under load)可能已有人在做 @evoludigit 今天认领。 未关闭
难度 1/5 1 小时以内 新手友好度 89/100
维护者通常 1 天内回复