PERFORMANCE.md states a 46% observer overhead that measures 0%, and the gate that should have caught it is one-sided by construction
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 65/100
- issue の種類
- バグ
- 明瞭さ
- おおむね明確
- 活発さ
- 活発
調査の方向性
Start by reading docs/PERFORMANCE.md:138-140, bench/check_regression.sh, and bench/baseline.txt, then rerun the observer benchmark with n=5 Ir measurements. Check the observer-gate history referenced by #915 and PR #1034. Done means the current cost is established, improvements cannot silently pass the gate, and the documented figures are mechanically checked or generated from the baseline.
索引モデルが issue の本文から書いたものです。
説明
Found by re-running a number the docs state in the present tense.
docs/PERFORMANCE.md:139 documents observer overhead as "~46% more (Ir
≈ 93.2M vs 63.8M)" and :138 as ~28% slower wall-clock.
Measured
Reported by Jon, then reproduced independently here (callgrind, this dev
box, current main build):
| documented | Jon | reproduced here | |
|---|---|---|---|
observed_loop Ir |
93,231,983 | 59,567,345 | 59,519,029 (−36.2%) |
unobserved_loop Ir |
63,843,716 | 59,573,261 | 59,525,213 (−6.8%) |
| observed − unobserved | +46% | −5,916 | −6,184 (−0.01%) |
| wall | 37 vs 29 ms | 6 vs 6 ms (median of 15) |
The observed loop now runs fewer instructions than the unobserved one.
That is noise around zero — not a win in the other direction — but the
documented cost of observation has gone from 46% to nothing, and the docs
still assert it in the present tense with a deterministic Ir figure
attached, which is the form most likely to be believed.
The two documented figures are byte-identical to bench/baseline.txt
(observed_loop 93231983, unobserved_loop 63843716), so this is one
staleness event surfacing in two places — the same shape as the
bench/baseline.txt staleness found this morning.
The gate cannot catch this, and that is the more useful half
bench/check_regression.sh compares one direction only:
inc=$(( (cur - base) * 100 / base ))
if [ "$inc" -gt "$THRESHOLD_PCT" ]; then # THRESHOLD_PCT=5
A drop of any size prints ok. observed_loop currently reads
-36% vs baseline and the gate is green.
So a one-sided gate cannot distinguish a real 36% improvement from a
36%-wrong baseline — both render identically, as good news. That is why
one unrecorded regeneration event desynced three artifacts at once and
nothing spoke: the baseline, the gate's effective threshold, and the
published docs.
docs/PERFORMANCE.md:140 compounds it by claiming the win is "now
regression-gated in both directions". It is gated in one.
What to do (measurement pass, not a quick edit)
- Establish what the observer costs now, properly: n=5 vs n=5, and
Irrather than wall (Irhere was stable; the wall figures are 6 ms
and unusable at this size). If it is genuinely ~0, find out when it
went to zero — the #915 observer gate (PR #1034) is the obvious
candidate and would meanunobserved:no longer buys anything on this
workload, which is a real documentation change, not a number swap. - Make the gate two-sided. An improvement beyond the threshold should
require the baseline to be re-pinned deliberately, exactly as a
regression does. A silent improvement is the signature of a stale
baseline far more often than of a win. - Make line 139 derived.
check_regression.sh --updateregenerates
the baseline; nothing regenerates the docs, so the copy in
PERFORMANCE.mdwill rot again on the next legitimate update. Either
generate that table frombench/baseline.txt, or enrol both figures in
tools/docs_claims_check.shso the doc fails when it disagrees with
the file it was copied from.
Note on what is NOT wrong here
The section immediately above (lines 100–120) is careful work: thresholds
placed from measured spreads, false-red rates counted across 14 and 5
runs, a --selftest that proves the check can fail. It was built to be
re-run. Lines 121–139 are a table from "a 2020-era x86-64 Linux dev box"
that can only be re-run by someone who thinks to.
That is the distinction worth keeping: prose lessons age well, measured
claims rot silently — and they rot in whatever copies were made of them.
Same point as the stale CLAUDE.md passage in ouroboros, now with a second
instance and a mechanical cause.
- 主要言語
- C
- スター
- 3
- フォーク
- 7
- 平均マージ
- 4時間 15分
- マージ済み PR(30日)
- 106
環境構築
このプロジェクトの開発コンテナを、あなたの GitHub アカウントでブラウザ上に起動します。
- Dockerfile または Docker Compose ファイルあり
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
InauguralSystems/EigenScript のほかの issue
-
area:lint-tooling bug
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
InauguralSystems/EigenScript#1340 ·
メンテナーはふだん 1 日以内に返信
-
area:stdlib found-by:code-review kind:silent-wrong
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
InauguralSystems/EigenScript#1338 ·
メンテナーはふだん 1 日以内に返信
-
area:lint-tooling found-by:critic kind:docs-drift
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
InauguralSystems/EigenScript#1335 ·
メンテナーはふだん 1 日以内に返信
-
area:ci found-by:critic kind:gate-defect
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
InauguralSystems/EigenScript#1311 ·
メンテナーはふだん 1 日以内に返信
-
enrolment: decide test_gc_runner_controls.py (exempt vs enrol) and whether floors need a ratchetオープンarea:gates found-by:critic kind:decision
難易度 2/5 1〜3時間 初心者へのやさしさ 65/100
InauguralSystems/EigenScript#1280 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
InauguralSystems/EigenScript の issue をすべて見る
似ている issue
-
bug needs triage
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
netdata/netdata#24062 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
-
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
BasedHardware/omi#19463 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
EchoTools/nevr-runtime#30 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
riscv-software-src/riscv-isa-sim#2448 ·
メンテナーはふだん 2 日以内に返信