A semicolon-separated .csv is indexed as one column

オープン
#42,353 コメント 0 件 リアクション 1 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
2/5
見積もり時間
1〜3時間
初心者へのやさしさ
45/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
pandas, python

調査の方向性

Start with api/core/rag/extractor/csv_extractor.py and inspect CSVExtractor.extract(), especially the pandas read_csv call and its csv_args. Reproduce the semicolon-separated example, then verify that CSV rows become separate named fields and add or run focused extractor tests for semicolon and tab-separated files.

索引モデルが issue の本文から書いたものです。

説明

Self Checks
  • I have read the Contributing Guide and Language Policy.
  • This is only for bug report, if you would like to ask a question, please head to Discussions.
  • I have searched for existing issues search for existing issues, including closed ones.
  • I confirm that I am using English to submit this report, otherwise it will be closed.
  • 【中文用户 & Non English User】请使用英语提交,否则会被关闭 :)
  • Please do not modify this template :) and fill in all the required fields.
Dify version

main (38f9d85d5a)

Cloud or Self Hosted

Self Hosted (Source)

Steps to reproduce

Upload a .csv file that is not comma separated to a knowledge base. This is the ordinary output of Save as → CSV in Excel on a machine whose locale uses a semicolon as the list separator — German, French, Spanish, Italian, and most of Europe — and it is also what many "export as CSV" buttons produce when the data contains commas.

data.csv:

name;region;units
widget;EU;12
gadget;US;7

api/core/rag/extractor/csv_extractor.py reads it with

df = pd.read_csv(csvfile, on_bad_lines="skip", **self.csv_args)

and self.csv_args carries no sep, so pandas uses its default, a comma.

Reproduced directly against the extractor:

>>> [d.page_content for d in CSVExtractor("data.csv", encoding="utf-8").extract()]
['name;region;units: widget;EU;12', 'name;region;units: gadget;US;7']
✔️ Expected Behavior

The columns of the file are the columns of the document:

['name: widget;region: EU;units: 12', 'name: gadget;region: US;units: 7']
❌ Actual Behavior

The file is read as having a single column. The whole row lands in it, separators included, and the one column is named name;region;units.

Nothing fails and nothing is logged — on_bad_lines="skip" is not reached, because a one-column parse is perfectly well formed. The chunk that reaches the index is the raw line with a column name glued to the front, so retrieval on any field of the sheet cannot match, and an answer built from it quotes the separators back to the user.

A tab separated export saved as .csv behaves the same way.

I have a fix and tests ready and will open a PR referencing this issue.

主要言語
TypeScript
スター
157k
フォーク
24.7k
平均マージ
22時間 32分
マージ済み PR(30日)
611

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

langgenius/dify のほかの issue

langgenius/dify の issue をすべて見る

似ている issue

TypeScript の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。