Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

CSVToDocument row mode drops a real extra_columns field on ragged rows

Open Beginner friendly
#12,994 0 comments 0 reactions 1 assignee View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
78/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python
Domain
data

Research direction

Start at the CSVToDocument entry point in haystack.components.converters.csv and trace row-mode handling for ragged CSV rows, including the rest key used by csv.DictReader. Add a regression test with a real extra_columns header and overflow values, then verify that the original field and overflow metadata are both preserved; compare with the ragged-row fix in #11944.

Written by the indexing model from the issue text.

Description

CSVToDocument(conversion_mode="row") uses extra_columns as the csv.DictReader rest key. If the CSV also has a column named extra_columns, a row with surplus fields overwrites that column before the converter sees it. The output document keeps the overflow values but silently loses the original field.

from haystack.components.converters.csv import CSVToDocument
from haystack.dataclasses import ByteStream

source = ByteStream(data=b"text,extra_columns\r\nhello,real value,overflow\r\n")
doc = CSVToDocument(conversion_mode="row").run(
    sources=[source], content_column="text"
)["documents"][0]
print(doc.content, doc.meta)
# hello {'extra_columns': "['overflow']", 'row_number': 0}

I ran this twice on current main; both runs produced the same result. The original real value cannot be recovered from the document. extra_columns is a valid CSV header, and an extra trailing delimiter or unquoted comma is enough to trigger this in row mode. File mode is unaffected.

Could the rest key use a name that cannot collide with a CSV header, or could the converter separate surplus fields before building metadata? A regression test should cover a real extra_columns column alongside overflow values. This is related to the ragged-row fix in #11944, but that fix only tests a header without this name. I can prepare the patch if the maintainers want one.

Dominant language
Python
Stars
26.6k
Forks
3.2k
Avg merge
1d 13h
Merged PRs (30d)
254

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from deepset-ai/haystack

All issues in deepset-ai/haystack

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.