A semicolon-separated .csv is indexed as one column
まだ誰も着手していません。
評価
- 難易度
- 2/5
- 見積もり時間
- 1〜3時間
- 初心者へのやさしさ
- 45/100
- issue の種類
- バグ
- 明瞭さ
- 明確に書かれている
- 活発さ
- 活発
- 技術スタック
- pandas, python
調査の方向性
Start with api/core/rag/extractor/csv_extractor.py and inspect CSVExtractor.extract(), especially the pandas read_csv call and its csv_args. Reproduce the semicolon-separated example, then verify that CSV rows become separate named fields and add or run focused extractor tests for semicolon and tab-separated files.
索引モデルが issue の本文から書いたものです。
説明
Self Checks
- I have read the Contributing Guide and Language Policy.
- This is only for bug report, if you would like to ask a question, please head to Discussions.
- I have searched for existing issues search for existing issues, including closed ones.
- I confirm that I am using English to submit this report, otherwise it will be closed.
- 【中文用户 & Non English User】请使用英语提交,否则会被关闭 :)
- Please do not modify this template :) and fill in all the required fields.
Dify version
main (38f9d85d5a)
Cloud or Self Hosted
Self Hosted (Source)
Steps to reproduce
Upload a .csv file that is not comma separated to a knowledge base. This is the ordinary output of Save as → CSV in Excel on a machine whose locale uses a semicolon as the list separator — German, French, Spanish, Italian, and most of Europe — and it is also what many "export as CSV" buttons produce when the data contains commas.
data.csv:
name;region;units
widget;EU;12
gadget;US;7
api/core/rag/extractor/csv_extractor.py reads it with
df = pd.read_csv(csvfile, on_bad_lines="skip", **self.csv_args)
and self.csv_args carries no sep, so pandas uses its default, a comma.
Reproduced directly against the extractor:
>>> [d.page_content for d in CSVExtractor("data.csv", encoding="utf-8").extract()]
['name;region;units: widget;EU;12', 'name;region;units: gadget;US;7']
✔️ Expected Behavior
The columns of the file are the columns of the document:
['name: widget;region: EU;units: 12', 'name: gadget;region: US;units: 7']
❌ Actual Behavior
The file is read as having a single column. The whole row lands in it, separators included, and the one column is named name;region;units.
Nothing fails and nothing is logged — on_bad_lines="skip" is not reached, because a one-column parse is perfectly well formed. The chunk that reaches the index is the raw line with a column name glued to the front, so retrieval on any field of the sheet cannot match, and an answer built from it quotes the separators back to the user.
A tab separated export saved as .csv behaves the same way.
I have a fix and tests ready and will open a PR referencing this issue.
- 主要言語
- TypeScript
- スター
- 157k
- フォーク
- 24.7k
- 平均マージ
- 22時間 32分
- マージ済み PR(30日)
- 611
コントリビューションガイド
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
langgenius/dify のほかの issue
-
Annotation Reply: a stored score threshold of 0.0 is silently replaced with 1, disabling the feature オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
langgenius/dify#42639 · コメント 1 件 · リアクション 1 件 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
langgenius/dify#42468 · コメント 1 件 · リアクション 1 件 ·
-
🐞 bug
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
langgenius/dify#42446 · リアクション 1 件 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
langgenius/dify#42355 · コメント 1 件 · リアクション 1 件 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
langgenius/dify#42350 · コメント 1 件 · リアクション 1 件 ·
langgenius/dify の issue をすべて見る
似ている issue
-
calcite-components needs triage refactor
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
Esri/calcite-design-system#15203 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 91/100
-
community first-timers-only good first issue hacktoberfest help wanted low hanging fruit up-for-grabs
難易度 1/5 1時間未満 初心者へのやさしさ 95/100
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
Automattic/studio#4908 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 90/100