Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

CSVToDocument row mode fails on CR-separated CSV input

Open Beginner friendly
#13,022 0 comments 0 reactions 1 assignee View on GitHub

Maintainers usually reply within 1 day

@julian-risch is already working on this.

Since Oct 1, 2026.

  • #13026 by @Marcuswang0824 — closed without merging

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
86/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python
Domain
data

Research direction

Start at the CSVToDocument conversion_mode="row" entry point and inspect how its csv.DictReader receives StringIO input. Run test_csv_to_document.py, then add or update the regression coverage for LF, CRLF, and CR records, quoted multiline fields, and row metadata. Done means all 20 tests pass and CR-separated input produces the expected documents.

Written by the indexing model from the issue text.

Description

P3

Describe the bug

CSVToDocument(conversion_mode="row") raises _csv.Error for CSV input whose records are separated by carriage returns (\r). The same input with LF or CRLF separators produces the expected documents. This was found with a local reproduction on current main, not a production incident.

To reproduce

from haystack.components.converters import CSVToDocument
from haystack.dataclasses import ByteStream

source = ByteStream(data=b"text,author\rfirst,Ada\rsecond,Bob\r")
result = CSVToDocument(conversion_mode="row").run(
    sources=[source], content_column="text"
)
print([doc.content for doc in result["documents"]])

Actual behavior

_csv.Error: new-line character seen in unquoted field - do you need to open the file with newline=''?

The error occurs while reading reader.fieldnames.

Expected behavior

The result should contain ['first', 'second'], with the corresponding author metadata and row numbers 0 and 1. Quoted multiline fields should retain their original newline characters.

Root cause and proposed fix

The converter passes io.StringIO(data) to csv.DictReader. Its default newline setting does not split CR-separated records. Passing io.StringIO(data, newline="") lets the CSV reader handle the line endings without translating quoted field content, consistent with Python's CSV input guidance.

A small local patch and regression test are ready. Before the fix, the parameterized LF/CRLF/CR test gives 1 failed / 2 passed; after the fix, all 20 tests in test_csv_to_document.py pass. The test also checks quoted multiline content and row metadata.

System

  • macOS arm64, Python 3.12.14
  • Haystack 3.3.0-rc0, main at bb5b39f42edb5cd9f69e5d9025d48d389c08598f
  • Official Hatch test environment; no model, GPU, or external API required

AI assistance: Codex generated this report, the reproduction, and the proposed patch, and ran the local checks.

Dominant language
Python
Stars
26.6k
Forks
3.2k
Avg merge
1d 21h
Merged PRs (30d)
257

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from deepset-ai/haystack

All issues in deepset-ai/haystack

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.