Unable to parse Document result in Python
Nobody has claimed this yet.
Assessment
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Newbie friendliness
- 35/100
Research direction
Start by reproducing the get_document_analysis response and inspect trp/init.py around Document._parse, Page._parse, and Line.init, as shown in the traceback. Done means the response parses without the reported KeyError and the affected block-reference case is covered by a test.
Written by the indexing model from the issue text.
Description
using textract-trp 0.1.3
When parsing "get_document_analysis" response the following output is generated:
Traceback (most recent call last):
File "G:\dev\OCR\main.py", line 17, in <module>
result = (textract.receive_document_result('52c4a450c667a18d89f4e26a1cf4b56859ad239f1a63279bec8f60458ae2284e'))
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "G:\dev\OCR\textract.py", line 62, in receive_document_result
return Document(response)
^^^^^^^^^^^^^^^^^^
File "G:\dev\OCR\venv\Lib\site-packages\trp\__init__.py", line 633, in __init__
self._parse()
File "G:\dev\OCR\venv\Lib\site-packages\trp\__init__.py", line 667, in _parse
page = Page(documentPage["Blocks"], self._blockMap)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "G:\dev\OCR\venv\Lib\site-packages\trp\__init__.py", line 516, in __init__
self._parse(blockMap)
File "G:\dev\OCR\venv\Lib\site-packages\trp\__init__.py", line 530, in _parse
l = Line(item, blockMap)
^^^^^^^^^^^^^^^^^^^^
File "G:\dev\OCR\venv\Lib\site-packages\trp\__init__.py", line 142, in __init__
if(blockMap[cid]["BlockType"] == "WORD"):
~~~~~~~~^^^^^
KeyError: '9e2f5e38-f865-4b79-a37b-ac8ed7a19f02'
- Dominant language
- Jupyter Notebook
- Stars
- 449
- Forks
- 262
- PR merge metrics
- No merged PRs in 30d
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from aws-samples/amazon-textract-code-samples
-
Difficulty 4/5 3-5 days Newbie friendliness 25/100
-
Difficulty 4/5 3-5 days Newbie friendliness 28/100
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 35/100
-
Difficulty 3/5 1-2 days Newbie friendliness 35/100
All issues in aws-samples/amazon-textract-code-samples
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
pymc-labs/pymc-marketing#3102 · 1 comment ·
Maintainers usually reply within 1 day
-
comp/plugins duplicate P3 platform/slack type/bug
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
NousResearch/hermes-agent#132594 · 2 comments ·
Maintainers usually reply within 1 day
-
needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
LucasSantana-Dev/Lucky#2637 ·
Maintainers usually reply within 1 day
-
bug needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
aws/aws-sdk-go-v2#3576 ·
Maintainers usually reply within 1 day
-
good first issue type: bug
Difficulty 1/5 Under an hour Newbie friendliness 88/100
medusajs/medusa#17127 · 2 comments ·
Maintainers usually reply within 1 day