Relative links in a page whose URL contains `//` are resolved with the slashes collapsed
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 3/5
- Thời gian dự kiến
- 1-2 ngày
- Mức phù hợp với người mới
- 68/100
Hướng nghiên cứu
Start in ArticleUrlRewriter.__call__, where the issue shows nested urllib.parse.urljoin calls for article_url, base_href, and item_url; read the resolver tests from #344 for the existing approach. Check both relative-link and base_href resolution with URLs containing consecutive slashes, and confirm the resulting ZIM path preserves the empty segments.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Follow-up to #344 / #340.
When a page's own URL contains consecutive slashes, relative links inside it are resolved with urllib.parse.urljoin, which drops the empty segments. So the // gets collapsed before normalize() even sees the URL.
Example (from @Sinkleberg's check in #344): rewriting other.html from https://example.com/x//y/page.html gives the ZIM path example.com/x/y/other.html, but a browser (WHATWG URL resolution) resolves it to https://example.com/x//y/other.html. So the link points to an entry that doesn't exist.
This is in ArticleUrlRewriter.__call__:
item_absolute_url = urljoin(
urljoin(self.article_url.value, base_href), item_url
)
The same applies to base_href resolution. A fix probably needs an RFC 3986 / WHATWG-style join that keeps empty segments instead of urljoin. #344 already has a small resolver like that in its tests.
- Ngôn ngữ chính
- Python
- Star
- 31
- Fork
- 31
- Merge trung bình
- 2 ngày 5 giờ
- Pull request đã merge (30 ngày)
- 3
Chuẩn bị môi trường
Dự án này không cung cấp dev container, Dockerfile hay hướng dẫn đóng góp, nên bạn cần tự thiết lập môi trường: hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của openzim/python-scraperlib
-
HTML rewriting: also rewrite `poster` attributeCó thể đã có người làm Có pull request liên kết đang mở hoặc đã được merge. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
openzim/python-scraperlib#339 · 1 bình luận ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
openzim/python-scraperlib#292 ·
-
URL normalisation: do not rewrite consecutive slashes `//` as a single slash `/`Có thể đã có người làm @anshuman83-40 đã nhận 7 ngày trước. Đang mở
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 45/100
openzim/python-scraperlib#340 ·
-
Add fuzzy rule to rewrite URLs of lesbases.anct.gouv.frCó thể làm lại được @benoit74 đã nhận 51 ngày trước và không có pull request nào đang mở. Đang mở
openzim/python-scraperlib#334 · 1 người được giao ·
-
Add utility to index ePub documents contentCó thể đã có người làm @Sriram-PR đã nhận 5 ngày trước. Đang mở
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 55/100
openzim/python-scraperlib#333 · 1 bình luận ·
Tất cả issue của openzim/python-scraperlib
Issue tương tự
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 60/100
521xueweihan/HelloGitHub#3924 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 67/100
wilbowes/EchoMuse#869 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 85/100
-
Claiming namespace `jft63`Đang mởnamespace operations
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 72/100
EclipseFdn/open-vsx.org#14043 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
test: TestServeUntilStale races the server's close against the client's sendall (BrokenPipeError under load)Có thể đã có người làm @evoludigit đã nhận hôm nay. Đang mở
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 89/100
Maintainer thường phản hồi trong vòng 1 ngày