CSVToDocument row mode fails on CR-separated CSV input
Maintainers usually reply within 1 day
@julian-risch is already working on this.
Since Oct 1, 2026.
- #13026 by @Marcuswang0824 — closed without merging
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 86/100
Research direction
Start at the CSVToDocument conversion_mode="row" entry point and inspect how its csv.DictReader receives StringIO input. Run test_csv_to_document.py, then add or update the regression coverage for LF, CRLF, and CR records, quoted multiline fields, and row metadata. Done means all 20 tests pass and CR-separated input produces the expected documents.
Written by the indexing model from the issue text.
Description
Describe the bug
CSVToDocument(conversion_mode="row") raises _csv.Error for CSV input whose records are separated by carriage returns (\r). The same input with LF or CRLF separators produces the expected documents. This was found with a local reproduction on current main, not a production incident.
To reproduce
from haystack.components.converters import CSVToDocument
from haystack.dataclasses import ByteStream
source = ByteStream(data=b"text,author\rfirst,Ada\rsecond,Bob\r")
result = CSVToDocument(conversion_mode="row").run(
sources=[source], content_column="text"
)
print([doc.content for doc in result["documents"]])
Actual behavior
_csv.Error: new-line character seen in unquoted field - do you need to open the file with newline=''?
The error occurs while reading reader.fieldnames.
Expected behavior
The result should contain ['first', 'second'], with the corresponding author metadata and row numbers 0 and 1. Quoted multiline fields should retain their original newline characters.
Root cause and proposed fix
The converter passes io.StringIO(data) to csv.DictReader. Its default newline setting does not split CR-separated records. Passing io.StringIO(data, newline="") lets the CSV reader handle the line endings without translating quoted field content, consistent with Python's CSV input guidance.
A small local patch and regression test are ready. Before the fix, the parameterized LF/CRLF/CR test gives 1 failed / 2 passed; after the fix, all 20 tests in test_csv_to_document.py pass. The test also checks quoted multiline content and row metadata.
System
- macOS arm64, Python 3.12.14
- Haystack 3.3.0-rc0,
mainatbb5b39f42edb5cd9f69e5d9025d48d389c08598f - Official Hatch test environment; no model, GPU, or external API required
AI assistance: Codex generated this report, the reproduction, and the proposed patch, and ran the local checks.
- Dominant language
- Python
- Stars
- 26.6k
- Forks
- 3.2k
- Avg merge
- 1d 21h
- Merged PRs (30d)
- 257
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from deepset-ai/haystack
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
deepset-ai/haystack#13199 ·
Maintainers usually reply within 1 day
-
CSVDocumentCleaner silently discards data for negative ignore countsPossibly taken @julian-risch claimed this 2 days ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
deepset-ai/haystack#13074 · 1 assignee ·
Maintainers usually reply within 1 day
-
CSVToDocument accepts an unsupported conversion_mode and silently converts in row modePossibly taken @Lesereingrape claimed this 7 days ago. OpenP3
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
deepset-ai/haystack#13070 ·
Maintainers usually reply within 1 day
-
PPTXToDocument drops soft line breaks when link_format is markdown or plainMay be free again A pull request for this issue was closed without being merged. OpenP3
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
deepset-ai/haystack#13043 · 1 comment ·
Maintainers usually reply within 1 day
-
InMemoryDocumentStore: count_unique_metadata_by_filter and get_metadata_field_unique_values disagree on int/float/bool valuesPossibly taken @julian-risch claimed this 8 days ago. OpenP3
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
deepset-ai/haystack#13040 · 1 comment · 1 assignee ·
Maintainers usually reply within 1 day
All issues in deepset-ai/haystack
Similar issues
-
request-theme
Difficulty 2/5 Under an hour Newbie friendliness 70/100
LizardByte/ThemerrDB#8877 · 1 comment ·
Maintainers usually reply within 1 day
-
area/install-update comp/gateway P0 sweeper:risk-compatibility type/bug
Difficulty 2/5 Under an hour Newbie friendliness 72/100
NousResearch/hermes-agent#135997 · 3 comments ·
Maintainers usually reply within 1 day
-
[BUG] JSONLoader rejects valid UTF-8 BOM filesPossibly taken @zouyonghe claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
anthropics/knowledge-work-plugins#1298 ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 90/100
chroma-core/chroma#7879 ·
Maintainers usually reply within 1 day