Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

--update runbook stamps the manifest from new_extraction AFTER build_merge() has mutated it in place (skills/*/references/update.md)

Aperta Adatta ai principianti
#2,865 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
2/5
Tempo stimato
1-3 ore
Idoneità per principianti
70/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Tranquilla
Stack tecnologico
python
Ambito
tooling

Direzione di ricerca

Apri skills/*/references/update.md e confronta il suo ordine con il percorso di estrazione in cli.py, dove _stamped_manifest_files(files_by_type, sem_result, ...) viene eseguito intorno a cli.py:3651 prima di _build_merge(...) intorno a cli.py:3881; il runbook ha build_merge a update.md:112 e lo stamp a update.md:153. Sposta la chiamata a _stamped_manifest_files sopra build_merge([new_extraction], ...) in modo che il manifest sia calcolato dall'estrazione non modificata, e aggiungi l'autocontrollo suggerito che stampa i file semantici timbrati rispetto ai file dispatchati che hanno prodotto nodi. Fatto significa che l'ordine del runbook rispecchia cli.py e un confronto sintetico prima/dopo di json.dumps(new_extraction) mostra che lo stamp viene ora letto prima del merge.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Summary

skills/*/references/update.md (the --update runbook) computes the manifest-stamp file list from new_extraction after passing it to build_merge(). build_merge() hands the extraction's node/edge dicts to build() → _fold_node_aliases / _coerce_non_string_ids / deduplicate_entities / build_from_json, which rewrite those dicts in place (ids re-keyed, source_file relativised, alias fields folded, edge source_file re-derived). Reading new_extraction again afterwards therefore no longer returns the run's own output, and _stamped_manifest_files(...) computed from it can be wrong.

cli.py's own extract path gets this right: _stamped_manifest_files(files_by_type, sem_result, ...) runs at ~cli.py:3651, before _build_merge(...) at ~cli.py:3881. The skill runbook has the opposite order (build_merge at update.md:112, _stamped_manifest_files at :153 on v8).

Observed (0.9.45, ~/Brain corpus, 2026-08-18)

Two incremental runs, same recipe, both merged correctly into graph.json but stamped the manifest wrong in both directions:

  • 14 changed docs, all extracted successfully → stamped=9 (5 legitimately-extracted files left with an empty semantic_hash, so they were re-queued and re-billed on the next run).
  • 10 changed docs → 18 stamp candidates, with 2 of the real 10 still left blank inside that inflated set.

A synthetic check confirms the mutation: json.dumps(new_extraction) before vs after build_merge([new_extraction], ...) differs (node source_file relativised, dedup rewrites applied). This is not a data-loss bug — the nodes are in the graph either way — but the manifest disagrees with the graph, which defeats the incremental gate.

Suggested fix

Move the stamp computation above the merge (one-line reorder), and print a self-check so a mismatch is visible:

new_extraction = json.loads(...)
incremental = json.loads(...)
from graphify.cli import _stamped_manifest_files
_manifest_files = _stamped_manifest_files(incremental['files'], new_extraction, Path('INPUT_PATH'))  # BEFORE build_merge
G = build_merge([new_extraction], ...)
...
save_manifest(_manifest_files, ...)
# stamped semantic files should equal (dispatched ∩ files that produced nodes/edges/hyperedges)

Alternatively build_merge() could copy.deepcopy its new_chunks argument, but that is a hidden cost on large extractions; fixing the runbook ordering matches what cli.py already does. Happy to open a PR for the runbook if useful.

Lingua principale
Python
Stelle
124k
Fork
11.9k
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Preparare l'ambiente

Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di Graphify-Labs/graphify

Tutte le issue di Graphify-Labs/graphify

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.