Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Feature: Post-RAG FactChecker Pipeline Component & Header-Aware Document Splitter

Aperta
#10,973 9 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
5/5
Tempo stimato
Più di una settimana
Idoneità per principianti
30/100
Tipo di issue
Funzionalità
Chiarezza
Da chiarire
Stato di attività
Tranquilla
Stack tecnologico
python
Ambito
ai, search

Direzione di ricerca

Inizia leggendo le API esistenti di Generator, DocumentSplitter, Pipeline e AnswerBuilder e le convenzioni dei relativi componenti. Chiarisci con i maintainer se FactChecker o HeaderAwareDocumentSplitter rientrano nell’ambito, poiché l’issue propone due componenti sostanziali. Il lavoro sarà considerato completato quando saranno disponibili un design concordato, unit test e documentazione per il componente selezionato.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

P2

Is your feature request related to a problem? Please describe.
A clear and concise description of what the problem is. Ex. I'm always frustrated when [...Yes. When building production RAG pipelines in Haystack, the pipeline typically ends with a Generator component (like OpenAIGenerator). While Haystack does a great job retrieving context, there is no native, deterministic post-processing component to audit the generated answer for hallucinations before returning it to the user.

Currently, if I want to guarantee that every sentence in the generated output is grounded in the retrieved Documents, I have to build a custom external loop to re-evaluate the output against the source context, which breaks the clean, linear flow of a Haystack Pipeline. Additionally, standard DocumentSplitter components often severe context by splitting paragraphs away from their Markdown headers strictly based on character/word counts.
]

Describe the solution you'd like
I would love to see two new components added to the Haystack ecosystem:

  1. FactChecker (or VerificationEvaluator) Component:
    A new Pipeline component designed to sit immediately after a Generator. It takes two inputs: the generated replies (String) and the documents (List[Document]). It uses a secondary LLM strictly as a judge to cross-reference the reply against the documents, effectively scoring it or stripping out unsupported claims. It would output a VerifiedReply and a ConfidenceScore.

  2. HeaderAwareDocumentSplitter Component:
    An enhancement or alternative to DocumentSplitter that parses Markdown/HTML AST. It refuses to split a chunk if it separates a heading (#, ##) from its immediate paragraph, ensuring that retrieved chunks retain their structural context.

Describe alternatives you've considered
I have considered using Haystack's AnswerBuilder with reference_pattern, but that simply matches regex citations, it doesn't deterministically verify if the LLM hallucinated the claim in the first place.

I have also considered using offline evaluation frameworks (like DeepEval or Ragas), but those are meant for testing datasets during development. I need a runtime component that actively blocks or warns users about unverified answers within the live production Pipeline. Currently, I am forced to write a Custom Component to handle this.

Additional context
I recently built a custom extraction pipeline in raw Python solving this exact problem, utilizing a multi-model approach (a fast 8B model for routing, and a 70B model for the final verification stage).

I am highly motivated to bring this pattern to Haystack. I am willing to write the PR, unit tests, and documentation for either the FactChecker component or the HeaderAwareDocumentSplitter if the core maintainers believe this aligns with Haystack's vision for production-ready RAG.

Lingua principale
Python
Stelle
26.6k
Fork
3.2k
Merge medio
1g 13h
PR unite (30g)
254

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di deepset-ai/haystack

Tutte le issue di deepset-ai/haystack

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.