A legacy .xls cell that says N/A or NULL is dropped from the index

Open
#42,415 1 comment 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
30/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Stale
Tech stack
pandas, python
Domain
backend, data

Research direction

Start with api/core/rag/extractor/excel_extractor.py and compare the pandas-based .xls branch with the openpyxl .xlsx branch. Review issue #42371 and its tests; done means literal values such as N/A and NULL remain available in indexed row output without dropping their column names.

Written by the indexing model from the issue text.

Description

Self Checks
  • I have read the Contributing Guide and Language Policy.
  • This is only for bug report, if you would like to ask a question, please head to Discussions.
  • I have searched for existing issues search for existing issues, including closed ones.
  • I confirm that I am using English to submit this report, otherwise it will be closed.
  • 【中文用户 & Non English User】请使用英语提交,否则会被关闭 :)
  • Please do not modify this template :) and fill in all the required fields.
Dify version

main (9c6c48b50b)

Cloud or Self Hosted

Self Hosted (Source)

Steps to reproduce

Upload a legacy .xls to a knowledge base in which some cells literally say N/A, NA, n/a, NULL, None, NaN or nan — "not applicable" in a hand-written sheet, NULL in a database export.

api/core/rag/extractor/excel_extractor.py reads the .xls branch through pandas:

df = excel_file.parse(sheet_name=sheet_name)

pandas treats those seven strings as missing values by default, so the cell arrives as NaN and is skipped together with its column name.

A three-row sheet, written with xlwt and read exactly as the extractor does:

Code | Status | Units        indexed today
A1   | N/A    | 12     ->    "Code":"A1";"Units":"12"
A2   | NULL   | 7            "Code":"A2";"Units":"7"
A3   | ok     | 3            "Code":"A3";"Status":"ok";"Units":"3"
✔️ Expected Behavior

A cell keeps the text it holds, as the .xlsx branch of the same extractor already does (it walks openpyxl cells and gets the string):

"Code":"A1";"Status":"N/A";"Units":"12"
"Code":"A2";"Status":"NULL";"Units":"7"
"Code":"A3";"Status":"ok";"Units":"3"
❌ Actual Behavior

The Status key is absent from the first two rows, not merely empty. A question such as "what is the status of A1?" finds no evidence in the index, and the chunk looks complete, so nothing signals the loss. dtype=object does not change this; only keep_default_na=False does.

A fix with tests is open in #42371.

Dominant language
TypeScript
Stars
157k
Forks
24.7k
Avg merge
22h 32m
Merged PRs (30d)
611

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from langgenius/dify

All issues in langgenius/dify

Similar issues

More TypeScript issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.