Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

--update runbook stamps the manifest from new_extraction AFTER build_merge() has mutated it in place (skills/*/references/update.md)

オープン 初心者向け
#2,865 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
2/5
見積もり時間
1〜3時間
初心者へのやさしさ
70/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
静か
技術スタック
python
領域
tooling

調査の方向性

skills/*/references/update.md を開き、その順序を cli.py の extract パスと比較してください。ここで _stamped_manifest_files(files_by_type, sem_result, ...) は cli.py:3651 付近で _build_merge(...) の前に実行され、_build_merge は cli.py:3881 付近にあります。runbook では build_merge が update.md:112 に、stamp が update.md:153 にあります。_stamped_manifest_files 呼び出しを build_merge([new_extraction], ...) の上に移動して、マニフェストが変更されていない抽出から計算されるようにし、提案されたセルフチェックを追加して、スタンプされたセマンティックファイルを、ノードを生成したディスパッチされたファイルと照合して出力するようにしてください。完了とは、runbook の順序が cli.py を反映し、json.dumps(new_extraction) の合成前後比較により、stamp がマージ前に読み取られるようになったことを示すことです。

索引モデルが issue の本文から書いたものです。

説明

Summary

skills/*/references/update.md (the --update runbook) computes the manifest-stamp file list from new_extraction after passing it to build_merge(). build_merge() hands the extraction's node/edge dicts to build() → _fold_node_aliases / _coerce_non_string_ids / deduplicate_entities / build_from_json, which rewrite those dicts in place (ids re-keyed, source_file relativised, alias fields folded, edge source_file re-derived). Reading new_extraction again afterwards therefore no longer returns the run's own output, and _stamped_manifest_files(...) computed from it can be wrong.

cli.py's own extract path gets this right: _stamped_manifest_files(files_by_type, sem_result, ...) runs at ~cli.py:3651, before _build_merge(...) at ~cli.py:3881. The skill runbook has the opposite order (build_merge at update.md:112, _stamped_manifest_files at :153 on v8).

Observed (0.9.45, ~/Brain corpus, 2026-08-18)

Two incremental runs, same recipe, both merged correctly into graph.json but stamped the manifest wrong in both directions:

  • 14 changed docs, all extracted successfully → stamped=9 (5 legitimately-extracted files left with an empty semantic_hash, so they were re-queued and re-billed on the next run).
  • 10 changed docs → 18 stamp candidates, with 2 of the real 10 still left blank inside that inflated set.

A synthetic check confirms the mutation: json.dumps(new_extraction) before vs after build_merge([new_extraction], ...) differs (node source_file relativised, dedup rewrites applied). This is not a data-loss bug — the nodes are in the graph either way — but the manifest disagrees with the graph, which defeats the incremental gate.

Suggested fix

Move the stamp computation above the merge (one-line reorder), and print a self-check so a mismatch is visible:

new_extraction = json.loads(...)
incremental = json.loads(...)
from graphify.cli import _stamped_manifest_files
_manifest_files = _stamped_manifest_files(incremental['files'], new_extraction, Path('INPUT_PATH'))  # BEFORE build_merge
G = build_merge([new_extraction], ...)
...
save_manifest(_manifest_files, ...)
# stamped semantic files should equal (dispatched ∩ files that produced nodes/edges/hyperedges)

Alternatively build_merge() could copy.deepcopy its new_chunks argument, but that is a hidden cost on large extractions; fixing the runbook ordering matches what cli.py already does. Happy to open a PR for the runbook if useful.

主要言語
Python
スター
124k
フォーク
11.9k
PR マージ指標
30日以内にマージされた PR はありません

環境構築

このプロジェクトには開発コンテナ、Dockerfile、コントリビューションガイドがありません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

Graphify-Labs/graphify のほかの issue

Graphify-Labs/graphify の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。