PERFORMANCE.md states a 46% observer overhead that measures 0%, and the gate that should have caught it is one-sided by construction
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 65/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Ambito
- build-system, documentation, performance, testing-qa
Direzione di ricerca
Start by reading docs/PERFORMANCE.md:138-140, bench/check_regression.sh, and bench/baseline.txt, then rerun the observer benchmark with n=5 Ir measurements. Check the observer-gate history referenced by #915 and PR #1034. Done means the current cost is established, improvements cannot silently pass the gate, and the documented figures are mechanically checked or generated from the baseline.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Found by re-running a number the docs state in the present tense.
docs/PERFORMANCE.md:139 documents observer overhead as "~46% more (Ir
≈ 93.2M vs 63.8M)" and :138 as ~28% slower wall-clock.
Measured
Reported by Jon, then reproduced independently here (callgrind, this dev
box, current main build):
| documented | Jon | reproduced here | |
|---|---|---|---|
observed_loop Ir |
93,231,983 | 59,567,345 | 59,519,029 (−36.2%) |
unobserved_loop Ir |
63,843,716 | 59,573,261 | 59,525,213 (−6.8%) |
| observed − unobserved | +46% | −5,916 | −6,184 (−0.01%) |
| wall | 37 vs 29 ms | 6 vs 6 ms (median of 15) |
The observed loop now runs fewer instructions than the unobserved one.
That is noise around zero — not a win in the other direction — but the
documented cost of observation has gone from 46% to nothing, and the docs
still assert it in the present tense with a deterministic Ir figure
attached, which is the form most likely to be believed.
The two documented figures are byte-identical to bench/baseline.txt
(observed_loop 93231983, unobserved_loop 63843716), so this is one
staleness event surfacing in two places — the same shape as the
bench/baseline.txt staleness found this morning.
The gate cannot catch this, and that is the more useful half
bench/check_regression.sh compares one direction only:
inc=$(( (cur - base) * 100 / base ))
if [ "$inc" -gt "$THRESHOLD_PCT" ]; then # THRESHOLD_PCT=5
A drop of any size prints ok. observed_loop currently reads
-36% vs baseline and the gate is green.
So a one-sided gate cannot distinguish a real 36% improvement from a
36%-wrong baseline — both render identically, as good news. That is why
one unrecorded regeneration event desynced three artifacts at once and
nothing spoke: the baseline, the gate's effective threshold, and the
published docs.
docs/PERFORMANCE.md:140 compounds it by claiming the win is "now
regression-gated in both directions". It is gated in one.
What to do (measurement pass, not a quick edit)
- Establish what the observer costs now, properly: n=5 vs n=5, and
Irrather than wall (Irhere was stable; the wall figures are 6 ms
and unusable at this size). If it is genuinely ~0, find out when it
went to zero — the #915 observer gate (PR #1034) is the obvious
candidate and would meanunobserved:no longer buys anything on this
workload, which is a real documentation change, not a number swap. - Make the gate two-sided. An improvement beyond the threshold should
require the baseline to be re-pinned deliberately, exactly as a
regression does. A silent improvement is the signature of a stale
baseline far more often than of a win. - Make line 139 derived.
check_regression.sh --updateregenerates
the baseline; nothing regenerates the docs, so the copy in
PERFORMANCE.mdwill rot again on the next legitimate update. Either
generate that table frombench/baseline.txt, or enrol both figures in
tools/docs_claims_check.shso the doc fails when it disagrees with
the file it was copied from.
Note on what is NOT wrong here
The section immediately above (lines 100–120) is careful work: thresholds
placed from measured spreads, false-red rates counted across 14 and 5
runs, a --selftest that proves the check can fail. It was built to be
re-run. Lines 121–139 are a table from "a 2020-era x86-64 Linux dev box"
that can only be re-run by someone who thinks to.
That is the distinction worth keeping: prose lessons age well, measured
claims rot silently — and they rot in whatever copies were made of them.
Same point as the stale CLAUDE.md passage in ouroboros, now with a second
instance and a mechanical cause.
- Lingua principale
- C
- Stelle
- 3
- Fork
- 7
- Merge medio
- 4h 1m
- PR unite (30g)
- 121
Preparare l'ambiente
Avvia il container di sviluppo del progetto nel browser, con il tuo account GitHub.
- Include un Dockerfile o un file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di InauguralSystems/EigenScript
-
area:ci kind:gate-defect
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
InauguralSystems/EigenScript#1448 ·
I maintainer di solito rispondono entro 1 giorno
-
area:ci kind:gate-defect
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
InauguralSystems/EigenScript#1431 ·
I maintainer di solito rispondono entro 1 giorno
-
area:gates kind:gate-defect
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
InauguralSystems/EigenScript#1429 ·
I maintainer di solito rispondono entro 1 giorno
-
area:docs kind:docs-drift
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 92/100
InauguralSystems/EigenScript#1428 ·
I maintainer di solito rispondono entro 1 giorno
-
area:docs good first issue kind:docs-drift
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 92/100
InauguralSystems/EigenScript#1400 ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di InauguralSystems/EigenScript
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
fastfetch-cli/fastfetch#2628 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
FujiNetWIFI/fujinet-firmware#1736 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
I maintainer di solito rispondono entro 3 giorni
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100