AutoMergingRetriever raises ValueError when a matched document has no parent (for example the root document)
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 82/100
Research direction
Start in haystack/components/retrievers/auto_merging_retriever.py and inspect _check_valid_documents alongside _try_merge_level. Reproduce the root-document case with the in-memory store, then add regression coverage showing parentless documents pass through while __level and __block_size validation remains enforced.
Written by the indexing model from the issue text.
Description
Describe the bug
AutoMergingRetriever.run() and run_async() call _check_valid_documents, which requires every matched document to have a truthy __parent_id meta field. When the matched set includes a document without a parent, such as the root document of the hierarchy or a document from outside it, the retriever raises ValueError: The matched leaf documents do not have the required meta field '__parent_id' and the whole pipeline fails.
This contradicts the retriever's own merge logic: _try_merge_level explicitly keeps documents that have no parents ("keep docs that have no parents"). That branch is unreachable today because validation rejects those documents first.
To Reproduce
Uses only public APIs and an in-memory document store; no model or external service is needed.
from haystack import Document
from haystack.components.preprocessors import HierarchicalDocumentSplitter
from haystack.components.retrievers.auto_merging_retriever import AutoMergingRetriever
from haystack.document_stores.in_memory import InMemoryDocumentStore
text = "The sun rose early in the morning. It cast a warm glow over the trees. Birds began to sing."
doc = Document(content=text)
splitter = HierarchicalDocumentSplitter(block_sizes={10, 3}, split_overlap=0, split_by="word")
docs = splitter.run([doc])["documents"]
# one store holds every level, so a retriever can match the root chunk itself
store = InMemoryDocumentStore()
store.write_documents(docs)
leaves = [d for d in docs if not d.meta["__children_ids"]]
root = [d for d in docs if d.meta["__level"] == 0][0]
retriever = AutoMergingRetriever(store, threshold=0.5)
print(retriever.run(leaves[:2])) # works
print(retriever.run(leaves[:2] + [root])) # raises ValueError
Actual output:
=== case 1: matched leaves only (baseline) ===
OK: [(2, 'The sun rose '), (2, 'early in the ')]
=== case 2: matched leaves + the root doc (one store holds every level) ===
ValueError: The matched leaf documents do not have the required meta field '__parent_id'
Expected behavior
Documents without __parent_id should pass through unchanged, as the existing pass-through branch in _try_merge_level intends, instead of failing the whole run.
Additional context
- In single-store setups the embedding retriever can legitimately return a parent or root chunk among the matches, so this input is realistic rather than off-contract.
- Distinct from #12943 and PR #12944: those fix merge ordering for leaves at different depths and do not touch
_check_valid_documents. - Proposed fix: drop the
__parent_idtruthiness requirement from_check_valid_documentsand keep the__level/__block_sizepresence checks, so parentless documents reach the existing pass-through branch. If strict validation is preferred, the error should at least name the offending document. I can open a PR with tests if either direction sounds right.
System
- OS: Linux x86_64
- Python: 3.14.7
- Haystack: current
main(3.3.0-rc0)
- Dominant language
- Python
- Stars
- 26.6k
- Forks
- 3.2k
- Avg merge
- 1d 13h
- Merged PRs (30d)
- 263
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from deepset-ai/haystack
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
deepset-ai/haystack#13022 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
deepset-ai/haystack#12994 · 1 assignee ·
Maintainers usually reply within 1 day
-
MarkdownHeaderSplitter treats headings inside longer closing fences as headersPossibly taken @julian-risch claimed this 3 days ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
deepset-ai/haystack#12954 · 2 comments · 1 assignee ·
Maintainers usually reply within 1 day
-
P3
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
deepset-ai/haystack#12945 ·
Maintainers usually reply within 1 day
-
Sockets.__getattribute__ references _sockets but the actual attribute is _sockets_dict — the optimization at lines 127-132 is dead code due to a typoPossibly taken @julian-risch claimed this 5 days ago. OpenP2
Difficulty 1/5 Under an hour Newbie friendliness 91/100
deepset-ai/haystack#12939 · 1 comment · 1 assignee ·
Maintainers usually reply within 1 day
All issues in deepset-ai/haystack
Similar issues
-
bug server
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
sportsdataverse/sportsdataverse-py#641 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
googleapis/google-cloud-python#18532 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day