Multipage pdf breaks if there is one blank page in between

Open
#19 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
25/100
Issue type
Bug
Clarity
Needs clarification
Activity status
Stale
Tech stack
aws, jupyter-notebook
Domain
backend, cloud

Research direction

The issue does not name a file, test, or entry point. Reproduce the failure with a multipage PDF containing a blank page, then locate the sample's PDF conversion flow and determine how blank or erroneous pages are handled. Done means the complete PDF conversion finishes without breaking on the intervening page.

Written by the indexing model from the issue text.

Description

Hello!
Thank you for such a wonderful library. We are using this extensively. We have one issue at hand. If we run a multipage pdf say of 200 pages and in between if any page is blank then it just breaks the complete pdf conversion.
Please suggest if there is a way we could avoid this so that the pdf gets converted by skipping the blank page or page with error.

Please guide.

Dominant language
Jupyter Notebook
Stars
449
Forks
262
PR merge metrics
No merged PRs in 30d

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from aws-samples/amazon-textract-code-samples

All issues in aws-samples/amazon-textract-code-samples

Similar issues

More Backend & API Design issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.