Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

[Bug]: One stray ")" in a page dictionary truncates numPages and makes all later pages unreachable.

Đang mở
#22,011 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
3/5
Thời gian dự kiến
1-2 ngày
Mức phù hợp với người mới
65/100
Loại issue
Lỗi
Độ rõ ràng
Đặc tả rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
javascript
Lĩnh vực
backend-api-design

Hướng nghiên cứu

The issue is in the lexer's handling of a stray ')', which triggers a FormatError in Lexer.getObj. The recovery path in Catalog.getAllPageDicts breaks early, truncating the page count. Start by examining src/core/lexer.js around case 0x29, then trace through src/core/catalog.js's getAllPageDicts method. The fix likely involves adjusting the error handling to not break the iteration on this specific error, or to treat it as a recoverable condition. Run the provided test PDF through the test suite to verify the fix restores numPages to 57.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

corrupted-pdf
Attach (recommended) or Link to PDF file

repro-57-pages-stray-paren.pdf

Web browser and its version

Chrome 153.0.0 (also reproduces in Node 23.11.1, so not browser-specific)

Operating system and its version

macOS 27.0 (26A428)

PDF.js version

5.4.296

Is the bug present in the latest PDF.js version?

Yes

Is a browser extension

No

Steps to reproduce the problem

The attached PDF declares 57 pages (/Count 57 in the root /Pages). Every page is well formed except page 40, whose dictionary contains a single stray ) before its >>.

  1. Open the attached file in the viewer.
  2. Read the page counter: it shows 40, not 57.
  3. Enter 41 in the page number box. The page cannot be reached.

Same result through the API, no viewer involved:

import { readFileSync } from "node:fs";
import { getDocument } from "pdfjs-dist/legacy/build/pdf.mjs";

const pdf = await getDocument({
  data: new Uint8Array(readFileSync("repro-57-pages-stray-paren.pdf")),
}).promise;

console.log(pdf.numPages);   // 40, the file declares 57
await pdf.getPage(39);       // ok
await pdf.getPage(40);       // UnknownErrorException: Illegal character: 41
await pdf.getPage(41);       // Error: Invalid page request.
What is the expected behavior?

numPages should be 57.

A page whose dictionary cannot be parsed may reasonably reject on its own getPage() call, but it should not remove the 17 well-formed pages that follow it from the document.

Two other implementations read all 57 pages from this file:

  • macOS PDFKit / Preview: mdls -name kMDItemNumberOfPages returns 57
  • Chrome's built-in viewer (PDFium): opens and renders all 57
What went wrong?

numPages is 40. Pages 1-39 load. Page 40 rejects with:

UnknownErrorException: Illegal character: 41
details: "FormatError: Illegal character: 41"

Pages 41 to 57 then reject with Error: Invalid page request. and stay unreachable for the lifetime of the document. 18 of 57 pages are lost.

Nothing in the public API signals this. getDocument() resolves successfully and numPages simply returns a smaller number, so a caller cannot tell a 40 page document apart from a 57 page document that silently lost 17 pages.

Link to a viewer

No response

Additional context

Where it appears to come from, reading the bundled worker (the code below is identical in 5.4.296 and 6.3.289):

  1. Lexer.getObj has case 0x29: which throws FormatError("Illegal character: " + ch). 0x29 is ). The neighbouring cases for >, { and } return a Cmd instead of throwing, so ) is the only delimiter in that switch that is fatal.
  2. PDFDocument.checkLastPage catches it and enters the recovery path Catalog.getAllPageDicts(recoveryMode).
  3. getAllPageDicts walks the page tree and breaks at the first kid that throws, then takes the page count from the partial map. That is where 57 becomes 40.

recoveryMode never engages here: every gate that enables it tests reason instanceof XRefEntryException, and a FormatError raised by the lexer is not one, so the truncated count stands.

Why it matters in practice: we hit this on a real 57 page document produced by a third-party authoring tool. Users were shown a 40 page document with no error of any kind, and had no way to know 17 pages were missing. The same file opens completely in Chrome and in Preview, so it appeared incomplete only where pdf.js was used.

The attached file is a minimal synthetic stand-in, 19 KB, with no third-party content. It produces the same result as the real document on 6.3.289 and on 5.4.296:

the real document attached file
declared /Count 57 57
pdf.js numPages 40 40
pages that load 1-39 1-39
page 40 Illegal character: 41 Illegal character: 41
pages 41-57 Invalid page request. Invalid page request.
Ngôn ngữ chính
JavaScript
Star
53.9k
Fork
10.7k
Merge trung bình
20 giờ 26 phút
Pull request đã merge (30 ngày)
122

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của mozilla/pdf.js

Tất cả issue của mozilla/pdf.js

Issue tương tự

Thêm issue về JavaScript

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.