--update runbook stamps the manifest from new_extraction AFTER build_merge() has mutated it in place (skills/*/references/update.md)
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
調査の方向性
skills/*/references/update.md を開き、その順序を cli.py の extract パスと比較してください。ここで _stamped_manifest_files(files_by_type, sem_result, ...) は cli.py:3651 付近で _build_merge(...) の前に実行され、_build_merge は cli.py:3881 付近にあります。runbook では build_merge が update.md:112 に、stamp が update.md:153 にあります。_stamped_manifest_files 呼び出しを build_merge([new_extraction], ...) の上に移動して、マニフェストが変更されていない抽出から計算されるようにし、提案されたセルフチェックを追加して、スタンプされたセマンティックファイルを、ノードを生成したディスパッチされたファイルと照合して出力するようにしてください。完了とは、runbook の順序が cli.py を反映し、json.dumps(new_extraction) の合成前後比較により、stamp がマージ前に読み取られるようになったことを示すことです。
索引モデルが issue の本文から書いたものです。
説明
Summary
skills/*/references/update.md (the --update runbook) computes the manifest-stamp file list from new_extraction after passing it to build_merge(). build_merge() hands the extraction's node/edge dicts to build() → _fold_node_aliases / _coerce_non_string_ids / deduplicate_entities / build_from_json, which rewrite those dicts in place (ids re-keyed, source_file relativised, alias fields folded, edge source_file re-derived). Reading new_extraction again afterwards therefore no longer returns the run's own output, and _stamped_manifest_files(...) computed from it can be wrong.
cli.py's own extract path gets this right: _stamped_manifest_files(files_by_type, sem_result, ...) runs at ~cli.py:3651, before _build_merge(...) at ~cli.py:3881. The skill runbook has the opposite order (build_merge at update.md:112, _stamped_manifest_files at :153 on v8).
Observed (0.9.45, ~/Brain corpus, 2026-08-18)
Two incremental runs, same recipe, both merged correctly into graph.json but stamped the manifest wrong in both directions:
- 14 changed docs, all extracted successfully →
stamped=9(5 legitimately-extracted files left with an emptysemantic_hash, so they were re-queued and re-billed on the next run). - 10 changed docs → 18 stamp candidates, with 2 of the real 10 still left blank inside that inflated set.
A synthetic check confirms the mutation: json.dumps(new_extraction) before vs after build_merge([new_extraction], ...) differs (node source_file relativised, dedup rewrites applied). This is not a data-loss bug — the nodes are in the graph either way — but the manifest disagrees with the graph, which defeats the incremental gate.
Suggested fix
Move the stamp computation above the merge (one-line reorder), and print a self-check so a mismatch is visible:
new_extraction = json.loads(...)
incremental = json.loads(...)
from graphify.cli import _stamped_manifest_files
_manifest_files = _stamped_manifest_files(incremental['files'], new_extraction, Path('INPUT_PATH')) # BEFORE build_merge
G = build_merge([new_extraction], ...)
...
save_manifest(_manifest_files, ...)
# stamped semantic files should equal (dispatched ∩ files that produced nodes/edges/hyperedges)
Alternatively build_merge() could copy.deepcopy its new_chunks argument, but that is a hidden cost on large extractions; fixing the runbook ordering matches what cli.py already does. Happy to open a PR for the runbook if useful.
- 主要言語
- Python
- スター
- 124k
- フォーク
- 11.9k
- PR マージ指標
- 30日以内にマージされた PR はありません
環境構築
このプロジェクトには開発コンテナ、Dockerfile、コントリビューションガイドがありません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
Graphify-Labs/graphify のほかの issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
Graphify-Labs/graphify#4241 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
Graphify-Labs/graphify#3763 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
Graphify-Labs/graphify#3611 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
Nix supportオープン
難易度 2/5 1〜3時間 初心者へのやさしさ 62/100
Graphify-Labs/graphify#3193 · リアクション 2 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 92/100
Graphify-Labs/graphify#2871 · コメント 2 件 ·
メンテナーはふだん 1 日以内に返信
Graphify-Labs/graphify の issue をすべて見る
似ている issue
-
documentation
難易度 1/5 1時間未満 初心者へのやさしさ 65/100
ansys/pydpf-core#3547 ·
メンテナーはふだん 1 日以内に返信
-
core
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
vectorize-io/hindsight#5457 ·
メンテナーはふだん 1 日以内に返信
-
[Bug]: LangChain drops OpenAI Responses text blocks from session recording対応中かも @ktz03 が今日担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
volcengine/OpenViking#5806 ·
メンテナーはふだん 1 日以内に返信
-
HTML: <template> content is extracted as document text対応中かも @ryanmeowy が今日担当しました。 オープンbug html
難易度 1/5 1時間未満 初心者へのやさしさ 82/100
docling-project/docling#4714 · コメント 2 件 ·
メンテナーはふだん 1 日以内に返信
-
APIv2 event data accepts a non-string reply and a NaN upper_bound対応中かも @awss1i が今日担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
freedomofpress/securedrop#7946 · コメント 1 件 ·
メンテナーはふだん 2 日以内に返信