Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Relative links in a page whose URL contains `//` are resolved with the slashes collapsed

未关闭
#346 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
3/5
预计耗时
1-2 天
新手友好度
68/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
活跃
技术栈
python
领域
backend

调研方向

Start in ArticleUrlRewriter.__call__, where the issue shows nested urllib.parse.urljoin calls for article_url, base_href, and item_url; read the resolver tests from #344 for the existing approach. Check both relative-link and base_href resolution with URLs containing consecutive slashes, and confirm the resulting ZIM path preserves the empty segments.

由索引模型根据 Issue 内容生成。

描述

Follow-up to #344 / #340.

When a page's own URL contains consecutive slashes, relative links inside it are resolved with urllib.parse.urljoin, which drops the empty segments. So the // gets collapsed before normalize() even sees the URL.

Example (from @Sinkleberg's check in #344): rewriting other.html from https://example.com/x//y/page.html gives the ZIM path example.com/x/y/other.html, but a browser (WHATWG URL resolution) resolves it to https://example.com/x//y/other.html. So the link points to an entry that doesn't exist.

This is in ArticleUrlRewriter.__call__:

item_absolute_url = urljoin(
    urljoin(self.article_url.value, base_href), item_url
)

The same applies to base_href resolution. A fix probably needs an RFC 3986 / WHATWG-style join that keeps empty segments instead of urljoin. #344 already has a small resolver like that in its tests.

主要语言
Python
星标
31
派生
31
平均合并
2 天 5 小时
30 天内合并 PR
3

环境准备

这个项目没有提供开发容器、Dockerfile 或贡献指南,环境需要你自己搭建:先看它的 README,通用步骤见我们的新手贡献指南。

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

openzim/python-scraperlib 的其他 Issue

查看 openzim/python-scraperlib 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。