--update runbook stamps the manifest from new_extraction AFTER build_merge() has mutated it in place (skills/*/references/update.md)
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 2/5
- Tiempo estimado
- 1-3 horas
- Aptitud para principiantes
- 70/100
Línea de trabajo
Abre skills/*/references/update.md y compara su orden con la ruta de extracción en cli.py, donde _stamped_manifest_files(files_by_type, sem_result, ...) se ejecuta alrededor de cli.py:3651 antes de _build_merge(...) alrededor de cli.py:3881; el runbook tiene build_merge en update.md:112 y el stamp en update.md:153. Mueve la llamada a _stamped_manifest_files por encima de build_merge([new_extraction], ...) para que el manifiesto se calcule a partir de la extracción sin modificar, y agrega la autocomprobación sugerida que imprime archivos semánticos sellados contra archivos despachados que produjeron nodos. Hecho significa que el orden del runbook refleja cli.py y una comparación sintética antes/después de json.dumps(new_extraction) muestra que el stamp ahora se lee antes del merge.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Summary
skills/*/references/update.md (the --update runbook) computes the manifest-stamp file list from new_extraction after passing it to build_merge(). build_merge() hands the extraction's node/edge dicts to build() → _fold_node_aliases / _coerce_non_string_ids / deduplicate_entities / build_from_json, which rewrite those dicts in place (ids re-keyed, source_file relativised, alias fields folded, edge source_file re-derived). Reading new_extraction again afterwards therefore no longer returns the run's own output, and _stamped_manifest_files(...) computed from it can be wrong.
cli.py's own extract path gets this right: _stamped_manifest_files(files_by_type, sem_result, ...) runs at ~cli.py:3651, before _build_merge(...) at ~cli.py:3881. The skill runbook has the opposite order (build_merge at update.md:112, _stamped_manifest_files at :153 on v8).
Observed (0.9.45, ~/Brain corpus, 2026-08-18)
Two incremental runs, same recipe, both merged correctly into graph.json but stamped the manifest wrong in both directions:
- 14 changed docs, all extracted successfully →
stamped=9(5 legitimately-extracted files left with an emptysemantic_hash, so they were re-queued and re-billed on the next run). - 10 changed docs → 18 stamp candidates, with 2 of the real 10 still left blank inside that inflated set.
A synthetic check confirms the mutation: json.dumps(new_extraction) before vs after build_merge([new_extraction], ...) differs (node source_file relativised, dedup rewrites applied). This is not a data-loss bug — the nodes are in the graph either way — but the manifest disagrees with the graph, which defeats the incremental gate.
Suggested fix
Move the stamp computation above the merge (one-line reorder), and print a self-check so a mismatch is visible:
new_extraction = json.loads(...)
incremental = json.loads(...)
from graphify.cli import _stamped_manifest_files
_manifest_files = _stamped_manifest_files(incremental['files'], new_extraction, Path('INPUT_PATH')) # BEFORE build_merge
G = build_merge([new_extraction], ...)
...
save_manifest(_manifest_files, ...)
# stamped semantic files should equal (dispatched ∩ files that produced nodes/edges/hyperedges)
Alternatively build_merge() could copy.deepcopy its new_chunks argument, but that is a hidden cost on large extractions; fixing the runbook ordering matches what cli.py already does. Happy to open a PR for the runbook if useful.
- Lenguaje dominante
- Python
- Estrellas
- 124k
- Forks
- 11.9k
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Preparar el entorno
Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de Graphify-Labs/graphify
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
Graphify-Labs/graphify#3763 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Graphify-Labs/graphify#3611 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Nix supportAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
Graphify-Labs/graphify#3193 · 2 reacciones ·
Los mantenedores suelen responder en 1 día
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 92/100
Graphify-Labs/graphify#2871 · 2 comentarios ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
Graphify-Labs/graphify#2840 ·
Los mantenedores suelen responder en 1 día
Todos los issues de Graphify-Labs/graphify
Issues similares
-
enhancement
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
epam/ai-dial-quickapps-backend#628 ·
Los mantenedores suelen responder en 2 días
-
0xlau.dev 已失效,切换成 timlau.meAbierto
Dificultad 1/5 Menos de una hora Aptitud para principiantes 69/100
timqian/chinese-independent-blogs#2235 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
eclipse-score/coverage_tool#27 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
bojieli/ai-agent-book#1174 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
RedHatQE/mtv-api-tests#721 ·
Los mantenedores suelen responder en 1 día