Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

--update runbook stamps the manifest from new_extraction AFTER build_merge() has mutated it in place (skills/*/references/update.md)

Abierto Apto para principiantes
#2,865 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
2/5
Tiempo estimado
1-3 horas
Aptitud para principiantes
70/100
Tipo de issue
Error
Claridad
Bien especificado
Estado de actividad
Tranquilo
Stack tecnológico
python
Área
tooling

Línea de trabajo

Abre skills/*/references/update.md y compara su orden con la ruta de extracción en cli.py, donde _stamped_manifest_files(files_by_type, sem_result, ...) se ejecuta alrededor de cli.py:3651 antes de _build_merge(...) alrededor de cli.py:3881; el runbook tiene build_merge en update.md:112 y el stamp en update.md:153. Mueve la llamada a _stamped_manifest_files por encima de build_merge([new_extraction], ...) para que el manifiesto se calcule a partir de la extracción sin modificar, y agrega la autocomprobación sugerida que imprime archivos semánticos sellados contra archivos despachados que produjeron nodos. Hecho significa que el orden del runbook refleja cli.py y una comparación sintética antes/después de json.dumps(new_extraction) muestra que el stamp ahora se lee antes del merge.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Summary

skills/*/references/update.md (the --update runbook) computes the manifest-stamp file list from new_extraction after passing it to build_merge(). build_merge() hands the extraction's node/edge dicts to build() → _fold_node_aliases / _coerce_non_string_ids / deduplicate_entities / build_from_json, which rewrite those dicts in place (ids re-keyed, source_file relativised, alias fields folded, edge source_file re-derived). Reading new_extraction again afterwards therefore no longer returns the run's own output, and _stamped_manifest_files(...) computed from it can be wrong.

cli.py's own extract path gets this right: _stamped_manifest_files(files_by_type, sem_result, ...) runs at ~cli.py:3651, before _build_merge(...) at ~cli.py:3881. The skill runbook has the opposite order (build_merge at update.md:112, _stamped_manifest_files at :153 on v8).

Observed (0.9.45, ~/Brain corpus, 2026-08-18)

Two incremental runs, same recipe, both merged correctly into graph.json but stamped the manifest wrong in both directions:

  • 14 changed docs, all extracted successfully → stamped=9 (5 legitimately-extracted files left with an empty semantic_hash, so they were re-queued and re-billed on the next run).
  • 10 changed docs → 18 stamp candidates, with 2 of the real 10 still left blank inside that inflated set.

A synthetic check confirms the mutation: json.dumps(new_extraction) before vs after build_merge([new_extraction], ...) differs (node source_file relativised, dedup rewrites applied). This is not a data-loss bug — the nodes are in the graph either way — but the manifest disagrees with the graph, which defeats the incremental gate.

Suggested fix

Move the stamp computation above the merge (one-line reorder), and print a self-check so a mismatch is visible:

new_extraction = json.loads(...)
incremental = json.loads(...)
from graphify.cli import _stamped_manifest_files
_manifest_files = _stamped_manifest_files(incremental['files'], new_extraction, Path('INPUT_PATH'))  # BEFORE build_merge
G = build_merge([new_extraction], ...)
...
save_manifest(_manifest_files, ...)
# stamped semantic files should equal (dispatched ∩ files that produced nodes/edges/hyperedges)

Alternatively build_merge() could copy.deepcopy its new_chunks argument, but that is a hidden cost on large extractions; fixing the runbook ordering matches what cli.py already does. Happy to open a PR for the runbook if useful.

Lenguaje dominante
Python
Estrellas
124k
Forks
11.9k
Métricas de merge de PR
Sin PR fusionados en 30 d

Preparar el entorno

Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de Graphify-Labs/graphify

Todos los issues de Graphify-Labs/graphify

Issues similares

Más issues de Python

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.