Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Memory-file extraction re-mints doc concepts under variant ids — near-duplicate nodes survive dedup

Aperta
#3,568 1 commento 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
55/100
Tipo di issue
Bug
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
python
Ambito
data

Direzione di ricerca

Start by tracing the merge path that combines extracted nodes, then inspect graph.json and the cached extraction entry under graphify-out/cache/semantic/…. Reproduce the memory-file update case and compare the canonical and variant concept ids. Done means repeated extraction does not leave near-duplicate concept nodes and the AMBIGUOUS edge is attached to the canonical node.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Summary

When a saved query memory (graphify-out/memory/query_*.md) is re-extracted on --update, concepts it mentions that already exist as doc-concept nodes can be minted under a different node id, producing a duplicate that id-based dedup cannot catch.

Evidence

Repo: rails-ai-bridge, graphifyy 0.9.61.

  • Original doc extraction of docs/port-registry-resolution.md minted docs_port_registry_resolution_resolver (label "Registry Resolver").
  • A memory file describing that same concept minted docs_port_registry_resolution_registry_resolver (label "Registry Resolver (port-registry-resolution doc concept)") — same stem, different entity suffix.
  • Both nodes coexisted in graph.json; the duplicate carried an AMBIGUOUS edge that belonged on the canonical node.

The node-id rule ({stem}_{entity}) makes the id deterministic from the label, so any label variation — parenthetical qualifiers, "Registry Resolver" vs "Resolver" — forks the id. Memory files quote concepts with extra qualifiers often enough that this recurs.

Suggested fix

One of:

  1. Label-normalized dedup for concept nodes: when merging, treat a new concept node as a duplicate if its normalized label (strip parentheticals, lowercase) matches an existing node with the same stem prefix — merge edges onto the canonical id.
  2. Canonical-id resolution at extraction: memory-file extraction could resolve mentioned concepts against the existing graph's node ids before minting new ones.

Option 1 is local to the merge path and doesn't require the extractor to know the existing graph.

Workaround applied locally

Rewrote the cached extraction entry (graphify-out/cache/semantic/…) to emit the canonical id and merged the node in graph.json. Holds until the memory file changes and the cache entry invalidates.

Lingua principale
Python
Stelle
124k
Fork
11.9k
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Preparare l'ambiente

Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di Graphify-Labs/graphify

Tutte le issue di Graphify-Labs/graphify

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.