Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

XLSXToDocument drops literal NA and N/A cell values

Aperta Adatta ai principianti
#12,945 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
2/5
Tempo stimato
1-3 ore
Idoneità per principianti
78/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
pandas, python
Ambito
data

Direzione di ricerca

Inizia dall’implementazione di XLSXToDocument e ispeziona la sua chiamata a pandas.read_excel e la gestione di read_excel_kwargs. Riproduci il caso del workbook, quindi aggiungi una copertura mirata per i valori letterali NA/N/A insieme alle celle realmente vuote; il lavoro è completato quando entrambi i valori letterali sopravvivono alla conversione mentre le celle vuote rimangono vuote.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

P3

XLSXToDocument turns text cells containing NA or N/A into empty output cells. These can be real values, such as a region code or an explicit status. Once the converter has emitted the Document, the original text cannot be recovered.

A small workbook with this sheet reproduces it:

Code Status
NA OK
N/A pending
from io import BytesIO
from openpyxl import Workbook
from haystack.components.converters.xlsx import XLSXToDocument
from haystack.dataclasses import ByteStream

book = Workbook()
sheet = book.active
sheet.append(["Code", "Status"])
sheet.append(["NA", "OK"])
sheet.append(["N/A", "pending"])
data = BytesIO()
book.save(data)

content = XLSXToDocument().run([ByteStream(data.getvalue())])["documents"][0].content
print(repr(content))
# ',A,B\n1,Code,Status\n2,,OK\n3,,pending\n'

I ran this twice on main; both outputs were identical. Markdown output also leaves the two cells blank. The converter uses pd.read_excel without overriding pandas' default NA-string parsing. Passing read_excel_kwargs={"keep_default_na": False} avoids the loss, but users have to know to opt out before converting the workbook.

Could the converter preserve literal text values by default while still keeping truly blank cells empty? A similar loss in the CSV preprocessing components is tracked in #12784, but this is the XLSX converter's separate read_excel path. I can prepare a focused patch and tests if you'd like one.

Lingua principale
Python
Stelle
26.6k
Fork
3.2k
Merge medio
1g 13h
PR unite (30g)
263

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di deepset-ai/haystack

Tutte le issue di deepset-ai/haystack

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.