--update runbook stamps the manifest from new_extraction AFTER build_merge() has mutated it in place (skills/*/references/update.md)
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 70/100
Research direction
Open skills/*/references/update.md and compare its order against the extract path in cli.py, where _stamped_manifest_files(files_by_type, sem_result, ...) runs around cli.py:3651 before _build_merge(...) around cli.py:3881; the runbook has build_merge at update.md:112 and the stamp at update.md:153. Move the _stamped_manifest_files call above build_merge([new_extraction], ...) so the manifest is computed from the unmutated extraction, and add the suggested self-check printing stamped semantic files against dispatched files that produced nodes. Done means the runbook order mirrors cli.py and a synthetic before/after json.dumps(new_extraction) comparison shows the stamp is now read pre-merge.
Written by the indexing model from the issue text.
Description
Summary
skills/*/references/update.md (the --update runbook) computes the manifest-stamp file list from new_extraction after passing it to build_merge(). build_merge() hands the extraction's node/edge dicts to build() → _fold_node_aliases / _coerce_non_string_ids / deduplicate_entities / build_from_json, which rewrite those dicts in place (ids re-keyed, source_file relativised, alias fields folded, edge source_file re-derived). Reading new_extraction again afterwards therefore no longer returns the run's own output, and _stamped_manifest_files(...) computed from it can be wrong.
cli.py's own extract path gets this right: _stamped_manifest_files(files_by_type, sem_result, ...) runs at ~cli.py:3651, before _build_merge(...) at ~cli.py:3881. The skill runbook has the opposite order (build_merge at update.md:112, _stamped_manifest_files at :153 on v8).
Observed (0.9.45, ~/Brain corpus, 2026-08-18)
Two incremental runs, same recipe, both merged correctly into graph.json but stamped the manifest wrong in both directions:
- 14 changed docs, all extracted successfully →
stamped=9(5 legitimately-extracted files left with an emptysemantic_hash, so they were re-queued and re-billed on the next run). - 10 changed docs → 18 stamp candidates, with 2 of the real 10 still left blank inside that inflated set.
A synthetic check confirms the mutation: json.dumps(new_extraction) before vs after build_merge([new_extraction], ...) differs (node source_file relativised, dedup rewrites applied). This is not a data-loss bug — the nodes are in the graph either way — but the manifest disagrees with the graph, which defeats the incremental gate.
Suggested fix
Move the stamp computation above the merge (one-line reorder), and print a self-check so a mismatch is visible:
new_extraction = json.loads(...)
incremental = json.loads(...)
from graphify.cli import _stamped_manifest_files
_manifest_files = _stamped_manifest_files(incremental['files'], new_extraction, Path('INPUT_PATH')) # BEFORE build_merge
G = build_merge([new_extraction], ...)
...
save_manifest(_manifest_files, ...)
# stamped semantic files should equal (dispatched ∩ files that produced nodes/edges/hyperedges)
Alternatively build_merge() could copy.deepcopy its new_chunks argument, but that is a hidden cost on large extractions; fixing the runbook ordering matches what cli.py already does. Happy to open a PR for the runbook if useful.
- Dominant language
- Python
- Stars
- 124k
- Forks
- 11.9k
- PR merge metrics
- No merged PRs in 30d
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Graphify-Labs/graphify
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Graphify-Labs/graphify#4241 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
Graphify-Labs/graphify#3763 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Graphify-Labs/graphify#3611 · 1 comment ·
Maintainers usually reply within 1 day
-
Nix supportOpen
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
Graphify-Labs/graphify#3193 · 2 reactions ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 92/100
Graphify-Labs/graphify#2871 · 2 comments ·
Maintainers usually reply within 1 day
All issues in Graphify-Labs/graphify
Similar issues
-
Difficulty 1/5 Under an hour Newbie friendliness 85/100
MystenLabs/MemWal#1163 · 2 comments ·
Maintainers usually reply within 1 day
-
infertopics leaves new nodes without a topic when untopiced neighbours outnumber topiced onesPossibly taken @moneebullah25 claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
FinanceFlash/unvibecode#218 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
NVIDIA/earth2studio#1241 ·
Maintainers usually reply within 3 days