benchmarks: code-intelligence result inaccessible; BENCHMARKS.md is a 404; headline metrics test conversation memory, not code graphs
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 42/100
- Tipo di issue
- Documentazione
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Ambito
- documentation, testing
Direzione di ricerca
Start with the README benchmarks table and its BENCHMARKS.md link, then check whether BENCHMARKS.md exists on main and review the references to #1706 and #2270. Confirm where the code-intelligence result and per-system tables are maintained. Done means restoring or correcting the link and documenting the requested INFERRED-edge precision result for the FastAPI demo corpus.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
The README benchmarks table shows three results — LOCOMO recall@10, LOCOMO QA accuracy, and LongMemEval-S QA accuracy — and directs to BENCHMARKS.md for "full per-system tables, the code-intelligence result, and reproduction commands." BENCHMARKS.md returns 404 on main.
Two problems follow from this:
1. The only code-relevant benchmark is unreachable
graphify's primary pitch is code knowledge graphs — AST parsing, cross-file link resolution, EXTRACTED vs INFERRED confidence tags. LOCOMO and LongMemEval benchmark conversational memory retrieval over chat transcripts. A user evaluating graphify for code navigation has no published number to look at. The code-intelligence result — presumably the metric that matters most for this tool's stated purpose — exists only in the missing file.
2. INFERRED edge precision for code graphs is unmeasured
The headline differentiator is the EXTRACTED / INFERRED tag: "you can tell what was read directly from what was inferred." The LOCOMO and LongMemEval results say nothing about whether INFERRED cross-file links in code are correct. Per #2270, confidence: EXTRACTED already appears on semantically-derived nodes; the complementary question — what fraction of INFERRED edges from AST resolution are accurate — has no number attached to it anywhere.
Suggested next steps
- Either commit BENCHMARKS.md to main or update the README link to wherever the code-intelligence result and per-system tables live. (#1706 notes the memory harness is a separate repo; the code-intelligence half presumably is too — a link would suffice.)
- Add a precision figure for INFERRED edges on the FastAPI demo corpus shown in the README. A 100-edge manual sample against ground-truth cross-file references would make the confidence-tag claim checkable without a full eval pipeline.
Happy to help with either.
- Lingua principale
- Python
- Stelle
- 124k
- Fork
- 11.9k
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Preparare l'ambiente
Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di Graphify-Labs/graphify
-
test(elixir): add a defguardp regression testForse già presa @ClockZW l’ha presa 1 giorno fa. Apertagood first issue help wanted
Difficoltà 2/5 1-3 ore Idoneità per principianti 90/100
Graphify-Labs/graphify#4076 ·
I maintainer di solito rispondono entro 1 giorno
-
test(php): parametrize the language-construct test across all constructsForse già presa @xiehuanyi l’ha presa oggi. Apertagood first issue help wanted
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
Graphify-Labs/graphify#4075 ·
I maintainer di solito rispondono entro 1 giorno
-
test(zig): assert a tagged-union nested-struct payload's fields are not mintedForse già presa @Jarvis-J-Jacob l’ha presa oggi. Apertagood first issue help wanted
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
Graphify-Labs/graphify#4074 ·
I maintainer di solito rispondono entro 1 giorno
-
test(rust): positive same-family cross-language base resolutionForse già presa @xiehuanyi l’ha presa oggi. Apertagood first issue help wanted
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
Graphify-Labs/graphify#4073 ·
I maintainer di solito rispondono entro 1 giorno
-
fix(astro): port the U+2028 trailing-comment terminator from the Svelte maskerForse già presa @Sourya-Prabaharan l’ha presa 1 giorno fa. Apertagood first issue help wanted
Difficoltà 2/5 1-3 ore Idoneità per principianti 90/100
Graphify-Labs/graphify#4072 ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di Graphify-Labs/graphify
Issue simili
-
needs-human needs-triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
gke-labs/kube-agents#2400 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Device Details tables: FS/SF columns contradict each other (nfet_01v8 Vt row, pfet_01v8 Idsat row)Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
google/skywater-pdk#450 ·
-
Drained trajectory arrays are overwritten when the sequence buffer is reusedForse già presa @sylvesterkaczmarek l’ha presa oggi. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
google-deepmind/bsuite#56 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
LearningCircuit/local-deep-research#7206 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
chingu-voyages/V62-tier3-team-33#285 ·
I maintainer di solito rispondono entro 1 giorno