Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

benchmarks: code-intelligence result inaccessible; BENCHMARKS.md is a 404; headline metrics test conversation memory, not code graphs

Aperta
#3,690 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
42/100
Tipo di issue
Documentazione
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
fastapi, python

Direzione di ricerca

Start with the README benchmarks table and its BENCHMARKS.md link, then check whether BENCHMARKS.md exists on main and review the references to #1706 and #2270. Confirm where the code-intelligence result and per-system tables are maintained. Done means restoring or correcting the link and documenting the requested INFERRED-edge precision result for the FastAPI demo corpus.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

help wanted

The README benchmarks table shows three results — LOCOMO recall@10, LOCOMO QA accuracy, and LongMemEval-S QA accuracy — and directs to BENCHMARKS.md for "full per-system tables, the code-intelligence result, and reproduction commands." BENCHMARKS.md returns 404 on main.

Two problems follow from this:

1. The only code-relevant benchmark is unreachable

graphify's primary pitch is code knowledge graphs — AST parsing, cross-file link resolution, EXTRACTED vs INFERRED confidence tags. LOCOMO and LongMemEval benchmark conversational memory retrieval over chat transcripts. A user evaluating graphify for code navigation has no published number to look at. The code-intelligence result — presumably the metric that matters most for this tool's stated purpose — exists only in the missing file.

2. INFERRED edge precision for code graphs is unmeasured

The headline differentiator is the EXTRACTED / INFERRED tag: "you can tell what was read directly from what was inferred." The LOCOMO and LongMemEval results say nothing about whether INFERRED cross-file links in code are correct. Per #2270, confidence: EXTRACTED already appears on semantically-derived nodes; the complementary question — what fraction of INFERRED edges from AST resolution are accurate — has no number attached to it anywhere.

Suggested next steps

  • Either commit BENCHMARKS.md to main or update the README link to wherever the code-intelligence result and per-system tables live. (#1706 notes the memory harness is a separate repo; the code-intelligence half presumably is too — a link would suffice.)
  • Add a precision figure for INFERRED edges on the FastAPI demo corpus shown in the README. A 100-edge manual sample against ground-truth cross-file references would make the confidence-tag claim checkable without a full eval pipeline.

Happy to help with either.

Lingua principale
Python
Stelle
124k
Fork
11.9k
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Preparare l'ambiente

Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di Graphify-Labs/graphify

Tutte le issue di Graphify-Labs/graphify

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.