--update runbook stamps the manifest from new_extraction AFTER build_merge() has mutated it in place (skills/*/references/update.md)
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 2/5
- Thời gian dự kiến
- 1-3 giờ
- Mức phù hợp với người mới
- 70/100
Hướng nghiên cứu
Mở skills/*/references/update.md và so sánh thứ tự của nó với đường dẫn extract trong cli.py, nơi _stamped_manifest_files(files_by_type, sem_result, ...) chạy quanh cli.py:3651 trước _build_merge(...) quanh cli.py:3881; runbook có build_merge tại update.md:112 và stamp tại update.md:153. Di chuyển lệnh gọi _stamped_manifest_files lên trên build_merge([new_extraction], ...) để manifest được tính từ extraction chưa bị thay đổi, và thêm self-check được đề xuất in các tệp ngữ nghĩa đã đóng dấu so với các tệp đã dispatch đã tạo ra các nút. Hoàn thành nghĩa là thứ tự runbook phản ánh cli.py và so sánh trước/sau tổng hợp của json.dumps(new_extraction) cho thấy stamp giờ được đọc trước khi merge.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Summary
skills/*/references/update.md (the --update runbook) computes the manifest-stamp file list from new_extraction after passing it to build_merge(). build_merge() hands the extraction's node/edge dicts to build() → _fold_node_aliases / _coerce_non_string_ids / deduplicate_entities / build_from_json, which rewrite those dicts in place (ids re-keyed, source_file relativised, alias fields folded, edge source_file re-derived). Reading new_extraction again afterwards therefore no longer returns the run's own output, and _stamped_manifest_files(...) computed from it can be wrong.
cli.py's own extract path gets this right: _stamped_manifest_files(files_by_type, sem_result, ...) runs at ~cli.py:3651, before _build_merge(...) at ~cli.py:3881. The skill runbook has the opposite order (build_merge at update.md:112, _stamped_manifest_files at :153 on v8).
Observed (0.9.45, ~/Brain corpus, 2026-08-18)
Two incremental runs, same recipe, both merged correctly into graph.json but stamped the manifest wrong in both directions:
- 14 changed docs, all extracted successfully →
stamped=9(5 legitimately-extracted files left with an emptysemantic_hash, so they were re-queued and re-billed on the next run). - 10 changed docs → 18 stamp candidates, with 2 of the real 10 still left blank inside that inflated set.
A synthetic check confirms the mutation: json.dumps(new_extraction) before vs after build_merge([new_extraction], ...) differs (node source_file relativised, dedup rewrites applied). This is not a data-loss bug — the nodes are in the graph either way — but the manifest disagrees with the graph, which defeats the incremental gate.
Suggested fix
Move the stamp computation above the merge (one-line reorder), and print a self-check so a mismatch is visible:
new_extraction = json.loads(...)
incremental = json.loads(...)
from graphify.cli import _stamped_manifest_files
_manifest_files = _stamped_manifest_files(incremental['files'], new_extraction, Path('INPUT_PATH')) # BEFORE build_merge
G = build_merge([new_extraction], ...)
...
save_manifest(_manifest_files, ...)
# stamped semantic files should equal (dispatched ∩ files that produced nodes/edges/hyperedges)
Alternatively build_merge() could copy.deepcopy its new_chunks argument, but that is a hidden cost on large extractions; fixing the runbook ordering matches what cli.py already does. Happy to open a PR for the runbook if useful.
- Ngôn ngữ chính
- Python
- Star
- 124k
- Fork
- 11.9k
- Chỉ số merge pull request
- Không có pull request nào được merge trong 30 ngày
Chuẩn bị môi trường
Dự án này không cung cấp dev container, Dockerfile hay hướng dẫn đóng góp, nên bạn cần tự thiết lập môi trường: hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của Graphify-Labs/graphify
-
[Bug]: `graphify export svg` writes a graph.svg that is not well-formed XML when a label contains a control characterCó thể đã có người làm Có pull request liên kết đang mở hoặc đã được merge. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
Graphify-Labs/graphify#4241 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
Graphify-Labs/graphify#3763 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Graphify-Labs/graphify#3611 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Nix supportĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 62/100
Graphify-Labs/graphify#3193 · 2 reaction ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
Graphify-Labs/graphify#2871 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của Graphify-Labs/graphify
Issue tương tự
-
first
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
AcademySoftwareFoundation/rmtc#54 · 1 bình luận ·
-
feature/cohorts feature/feature-flags team/feature-flags
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
Maintainer thường phản hồi trong vòng 1 ngày
-
License examples/ as MITCó thể đã có người làm @PGrayCS đã nhận hôm nay. Đang mởdocumentation enhancement example good first issue
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
speedyk-005/yasbd-lib#383 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
interactions-py/interactions.py#1827 ·
-
Managed start can fail when OpenVMM reads its control capability before NVX writes itCó thể đã có người làm @ppenna đã nhận hôm nay. Đang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
Maintainer thường phản hồi trong vòng 1 ngày