A semicolon-separated .csv is indexed as one column

未关闭
#42,353 0 条评论 1 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
2/5
预计耗时
1-3 小时
新手友好度
45/100
Issue 类型
缺陷
描述清晰度
描述清楚
活跃度
活跃
技术栈
pandas, python

调研方向

Start with api/core/rag/extractor/csv_extractor.py and inspect CSVExtractor.extract(), especially the pandas read_csv call and its csv_args. Reproduce the semicolon-separated example, then verify that CSV rows become separate named fields and add or run focused extractor tests for semicolon and tab-separated files.

由索引模型根据 Issue 内容生成。

描述

Self Checks
  • I have read the Contributing Guide and Language Policy.
  • This is only for bug report, if you would like to ask a question, please head to Discussions.
  • I have searched for existing issues search for existing issues, including closed ones.
  • I confirm that I am using English to submit this report, otherwise it will be closed.
  • 【中文用户 & Non English User】请使用英语提交,否则会被关闭 :)
  • Please do not modify this template :) and fill in all the required fields.
Dify version

main (38f9d85d5a)

Cloud or Self Hosted

Self Hosted (Source)

Steps to reproduce

Upload a .csv file that is not comma separated to a knowledge base. This is the ordinary output of Save as → CSV in Excel on a machine whose locale uses a semicolon as the list separator — German, French, Spanish, Italian, and most of Europe — and it is also what many "export as CSV" buttons produce when the data contains commas.

data.csv:

name;region;units
widget;EU;12
gadget;US;7

api/core/rag/extractor/csv_extractor.py reads it with

df = pd.read_csv(csvfile, on_bad_lines="skip", **self.csv_args)

and self.csv_args carries no sep, so pandas uses its default, a comma.

Reproduced directly against the extractor:

>>> [d.page_content for d in CSVExtractor("data.csv", encoding="utf-8").extract()]
['name;region;units: widget;EU;12', 'name;region;units: gadget;US;7']
✔️ Expected Behavior

The columns of the file are the columns of the document:

['name: widget;region: EU;units: 12', 'name: gadget;region: US;units: 7']
❌ Actual Behavior

The file is read as having a single column. The whole row lands in it, separators included, and the one column is named name;region;units.

Nothing fails and nothing is logged — on_bad_lines="skip" is not reached, because a one-column parse is perfectly well formed. The chunk that reaches the index is the raw line with a column name glued to the front, so retrieval on any field of the sheet cannot match, and an answer built from it quotes the separators back to the user.

A tab separated export saved as .csv behaves the same way.

I have a fix and tests ready and will open a PR referencing this issue.

主要语言
TypeScript
星标
157k
派生
24.7k
平均合并
22 小时 32 分钟
30 天内合并 PR
611

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

langgenius/dify 的其他 Issue

查看 langgenius/dify 的全部 Issue

相似的 Issue

更多 TypeScript Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。