Topological overfitting detection: H0 gap catches overfitting before accuracy diverges (r=0.998)
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 15/100
- issue の種類
- 機能追加
- 明瞭さ
- 説明が足りない
- 活発さ
- 静か
- 技術スタック
- python
調査の方向性
この issue では、MLCommons Algorithmic Efficiency リポジトリ内のファイル、テスト、エントリーポイントが指定されていません。まず、リポジトリのベンチマークへの貢献に関するガイドラインを見つけ、外部の ph-training パイプラインと ripser ベースの手法がその対象範囲に適合するかを判断してください。完了とするには、定義された統合ポイント、再現可能なベンチマークの証拠、および受け入れられた検証基準が必要です。
索引モデルが issue の本文から書いたものです。
説明
Summary
We found that Persistent Homology (H0 total persistence) on class-mean direction vectors provides a real-time overfitting signal with r=0.998 correlation to the generalization gap — often detecting overfitting before the train/test accuracy gap becomes visible.
Method
- Extract direction vectors from model:
d = normalize(engine_A(x) - engine_G(x)) - Compute per-class mean directions
- Build cosine distance matrix between class centroids
- Run H0 persistent homology (via ripser)
- Compare H0_train vs H0_test — the gap predicts overfitting
Also includes
- Automatic LR search: The LR that minimizes H0 CV (coefficient of variation) over 5 epochs = optimal LR
- 1-epoch difficulty prediction: H0 after 1 epoch predicts final accuracy (H0=4.38 → 98.3%, H0=2.02 → 52.0%)
- Confusion prediction: H0 merge order = confusion pairs (Spearman r=-0.97)
Verified results
| Dataset | Accuracy | Best LR | Early Stop | Time |
|---|---|---|---|---|
| MNIST | 98.3% | 1e-03 | no | 2.2 min |
| Fashion | 87.4% | 3e-04 | no | 2.2 min |
| CIFAR-10 | 52.0% | 1e-03 | yes (ep 6) | 1.4 min |
CIFAR early-stopped at epoch 6 when H0_gap exceeded threshold — preventing wasted compute on a model that was already overfitting.
Repo: https://github.com/need-singularity/ph-training
Install: pip install -e . then ph-train --dataset cifar
Related projects
- logout — Consciousness Continuity Engine. The main research project with the dual-engine (PureFieldEngine) architecture that produces direction vectors analyzed by PH.
- Anima — Conversational consciousness agent with real-time PH overfitting detection integrated into the live inference loop.
- ph-training — Standalone training pipeline.
pip install -e .thenph-train --dataset cifar.
- 主要言語
- Python
- スター
- 425
- フォーク
- 78
- PR マージ指標
- 30日以内にマージされた PR はありません
環境構築
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
mlcommons/algorithmic-efficiency のほかの issue
-
難易度 5/5 1週間以上 初心者へのやさしさ 30/100
-
難易度 4/5 3〜5日 初心者へのやさしさ 38/100
mlcommons/algorithmic-efficiency#924 · コメント 5 件 ·
-
Anima: live PH monitoring in conversational agent — overfitting detection every 50 interactionsオープン
難易度 5/5 1週間以上 初心者へのやさしさ 25/100
-
難易度 5/5 1週間以上 初心者へのやさしさ 20/100
-
難易度 3/5 1〜2日 初心者へのやさしさ 48/100
mlcommons/algorithmic-efficiency の issue をすべて見る
似ている issue
-
ACK_WAITING HELP_WANTED UPDATE_CS
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
OWASP/CheatSheetSeries#2458 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 90/100
BasedHardware/omi#19711 ·
メンテナーはふだん 1 日以内に返信
-
Qwen3_5MoeModel no longer returns router_logits, breaking aux loss with output_router_logits=Trueオープン
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
huggingface/transformers#49172 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
vllm-project/vllm-metal#885 ·
メンテナーはふだん 1 日以内に返信