bug: pdf splitting modifies returned csv elements

Open Beginner friendly
#201 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
62/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Stale
Tech stack
python
Domain
api

Research direction

Start with _test_unstructured_client/integration/test_decorators.py::test_integration_split_csv_response and replace its shortened assertion with a comparison that exposes the extra text_as_html field and differing element id. Trace the CSV response handling for split_pdf_page=True and False; done means both responses are identical and the integration test checks the full result.

Written by the indexing model from the issue text.

Description

bug

Describe the bug
When specifying output_format as csv, the response from the api is different when split_pdf_page is True or False. When False, the elements contain an extra metadata field: text_as_html. This also means the element id does not match.

To Reproduce
_test_unstructured_client/integration/test_decorators.py::test_integration_split_csv_response illustrates this, but is passing because it asserts on a shortened string.

Expected behavior
The response to be identical whether or not split_pdf_page is True or False.

Dominant language
Python
Stars
119
Forks
22
Avg merge
1d 21h
Merged PRs (30d)
1

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Unstructured-IO/unstructured-python-client

All issues in Unstructured-IO/unstructured-python-client

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.