Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

DOCXToDocument drops hyperlink addresses inside tables

Aperta
#12,977 2 commenti 0 reazioni 1 assegnatario Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
3/5
Tempo stimato
1-2 giorni
Idoneità per principianti
78/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
python
Ambito
backend

Direzione di ricerca

Start with DOCXToDocument and trace _extract_elements(), _process_links_in_paragraph(), _table_to_markdown(), and _table_to_csv(). Reproduce the issue with hyperlinks in a body paragraph and table cell using markdown, plain, and default link formats. Add tests covering both table formats and link formats; done means table-cell addresses are preserved except when link_format is none.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

P3

Describe the bug
DOCXToDocument(link_format="markdown") preserves a hyperlink in a body paragraph but drops the address when the same hyperlink is inside a table cell. This loses source links from tables during indexing. link_format="plain" has the same problem.

Error message
None. Conversion succeeds with incomplete content.

Expected behavior
Links in table cells should use the chosen link_format, as links in body paragraphs do. The default link_format="none" should remain unchanged.

To reproduce
Create a DOCX with a body paragraph and a table cell, each containing a hyperlink with display text docs and address https://example.com/reference. Run:

from haystack.components.converters.docx import DOCXToDocument

content = DOCXToDocument(link_format="markdown", table_format="markdown").run(
    sources=["links.docx"]
)["documents"][0].content
print(content)

On current main (8a5406e), two identical runs returned:

Outside [docs](https://example.com/reference)
| Inside docs |
| ----------- |

The first link survives; the table link does not. _extract_elements() calls _process_links_in_paragraph() for body paragraphs, while _table_to_markdown() and _table_to_csv() read cell.text, which has no link address. A small fix would format links in cell paragraphs before serializing either table format, with tests for markdown/plain links and the unchanged default.

System

  • OS: Linux
  • Haystack checkout: main at 8a5406e (local package reports 3.2.0)
  • python-docx: 1.2.0
Lingua principale
Python
Stelle
26.6k
Fork
3.2k
Merge medio
1g 13h
PR unite (30g)
241

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di deepset-ai/haystack

Tutte le issue di deepset-ai/haystack

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.