A semicolon-separated .csv is indexed as one column
还没有人认领这个 Issue。
评估
- 难度
- 2/5
- 预计耗时
- 1-3 小时
- 新手友好度
- 45/100
- Issue 类型
- 缺陷
- 描述清晰度
- 描述清楚
- 活跃度
- 活跃
- 技术栈
- pandas, python
调研方向
Start with api/core/rag/extractor/csv_extractor.py and inspect CSVExtractor.extract(), especially the pandas read_csv call and its csv_args. Reproduce the semicolon-separated example, then verify that CSV rows become separate named fields and add or run focused extractor tests for semicolon and tab-separated files.
由索引模型根据 Issue 内容生成。
描述
Self Checks
- I have read the Contributing Guide and Language Policy.
- This is only for bug report, if you would like to ask a question, please head to Discussions.
- I have searched for existing issues search for existing issues, including closed ones.
- I confirm that I am using English to submit this report, otherwise it will be closed.
- 【中文用户 & Non English User】请使用英语提交,否则会被关闭 :)
- Please do not modify this template :) and fill in all the required fields.
Dify version
main (38f9d85d5a)
Cloud or Self Hosted
Self Hosted (Source)
Steps to reproduce
Upload a .csv file that is not comma separated to a knowledge base. This is the ordinary output of Save as → CSV in Excel on a machine whose locale uses a semicolon as the list separator — German, French, Spanish, Italian, and most of Europe — and it is also what many "export as CSV" buttons produce when the data contains commas.
data.csv:
name;region;units
widget;EU;12
gadget;US;7
api/core/rag/extractor/csv_extractor.py reads it with
df = pd.read_csv(csvfile, on_bad_lines="skip", **self.csv_args)
and self.csv_args carries no sep, so pandas uses its default, a comma.
Reproduced directly against the extractor:
>>> [d.page_content for d in CSVExtractor("data.csv", encoding="utf-8").extract()]
['name;region;units: widget;EU;12', 'name;region;units: gadget;US;7']
✔️ Expected Behavior
The columns of the file are the columns of the document:
['name: widget;region: EU;units: 12', 'name: gadget;region: US;units: 7']
❌ Actual Behavior
The file is read as having a single column. The whole row lands in it, separators included, and the one column is named name;region;units.
Nothing fails and nothing is logged — on_bad_lines="skip" is not reached, because a one-column parse is perfectly well formed. The chunk that reaches the index is the raw line with a column name glued to the front, so retrieval on any field of the sheet cannot match, and an answer built from it quotes the separators back to the user.
A tab separated export saved as .csv behaves the same way.
I have a fix and tests ready and will open a PR referencing this issue.
- 主要语言
- TypeScript
- 星标
- 157k
- 派生
- 24.7k
- 平均合并
- 22 小时 32 分钟
- 30 天内合并 PR
- 611
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
langgenius/dify 的其他 Issue
-
Annotation Reply: a stored score threshold of 0.0 is silently replaced with 1, disabling the feature 未关闭
难度 2/5 1-3 小时 新手友好度 88/100
langgenius/dify#42639 · 1 条评论 · 1 个 reaction ·
-
难度 2/5 1-3 小时 新手友好度 70/100
langgenius/dify#42468 · 1 条评论 · 1 个 reaction ·
-
🐞 bug
难度 2/5 1-3 小时 新手友好度 86/100
langgenius/dify#42446 · 1 个 reaction ·
-
难度 2/5 1-3 小时 新手友好度 88/100
langgenius/dify#42355 · 1 条评论 · 1 个 reaction ·
-
难度 2/5 1-3 小时 新手友好度 88/100
langgenius/dify#42350 · 1 条评论 · 1 个 reaction ·
相似的 Issue
-
calcite-components needs triage refactor
难度 2/5 1-3 小时 新手友好度 75/100
Esri/calcite-design-system#15203 ·
-
难度 2/5 1-3 小时 新手友好度 91/100
-
community first-timers-only good first issue hacktoberfest help wanted low hanging fruit up-for-grabs
难度 1/5 1 小时以内 新手友好度 95/100
-
难度 2/5 1-3 小时 新手友好度 78/100
Automattic/studio#4908 ·
-
难度 2/5 1-3 小时 新手友好度 90/100