Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

--update runbook stamps the manifest from new_extraction AFTER build_merge() has mutated it in place (skills/*/references/update.md)

Open Beginner friendly
#2,865 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
70/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Quiet
Tech stack
python
Domain
tooling

Research direction

Open skills/*/references/update.md and compare its order against the extract path in cli.py, where _stamped_manifest_files(files_by_type, sem_result, ...) runs around cli.py:3651 before _build_merge(...) around cli.py:3881; the runbook has build_merge at update.md:112 and the stamp at update.md:153. Move the _stamped_manifest_files call above build_merge([new_extraction], ...) so the manifest is computed from the unmutated extraction, and add the suggested self-check printing stamped semantic files against dispatched files that produced nodes. Done means the runbook order mirrors cli.py and a synthetic before/after json.dumps(new_extraction) comparison shows the stamp is now read pre-merge.

Written by the indexing model from the issue text.

Description

Summary

skills/*/references/update.md (the --update runbook) computes the manifest-stamp file list from new_extraction after passing it to build_merge(). build_merge() hands the extraction's node/edge dicts to build() → _fold_node_aliases / _coerce_non_string_ids / deduplicate_entities / build_from_json, which rewrite those dicts in place (ids re-keyed, source_file relativised, alias fields folded, edge source_file re-derived). Reading new_extraction again afterwards therefore no longer returns the run's own output, and _stamped_manifest_files(...) computed from it can be wrong.

cli.py's own extract path gets this right: _stamped_manifest_files(files_by_type, sem_result, ...) runs at ~cli.py:3651, before _build_merge(...) at ~cli.py:3881. The skill runbook has the opposite order (build_merge at update.md:112, _stamped_manifest_files at :153 on v8).

Observed (0.9.45, ~/Brain corpus, 2026-08-18)

Two incremental runs, same recipe, both merged correctly into graph.json but stamped the manifest wrong in both directions:

  • 14 changed docs, all extracted successfully → stamped=9 (5 legitimately-extracted files left with an empty semantic_hash, so they were re-queued and re-billed on the next run).
  • 10 changed docs → 18 stamp candidates, with 2 of the real 10 still left blank inside that inflated set.

A synthetic check confirms the mutation: json.dumps(new_extraction) before vs after build_merge([new_extraction], ...) differs (node source_file relativised, dedup rewrites applied). This is not a data-loss bug — the nodes are in the graph either way — but the manifest disagrees with the graph, which defeats the incremental gate.

Suggested fix

Move the stamp computation above the merge (one-line reorder), and print a self-check so a mismatch is visible:

new_extraction = json.loads(...)
incremental = json.loads(...)
from graphify.cli import _stamped_manifest_files
_manifest_files = _stamped_manifest_files(incremental['files'], new_extraction, Path('INPUT_PATH'))  # BEFORE build_merge
G = build_merge([new_extraction], ...)
...
save_manifest(_manifest_files, ...)
# stamped semantic files should equal (dispatched ∩ files that produced nodes/edges/hyperedges)

Alternatively build_merge() could copy.deepcopy its new_chunks argument, but that is a hidden cost on large extractions; fixing the runbook ordering matches what cli.py already does. Happy to open a PR for the runbook if useful.

Dominant language
Python
Stars
124k
Forks
11.9k
PR merge metrics
No merged PRs in 30d

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Graphify-Labs/graphify

All issues in Graphify-Labs/graphify

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.