Epic: ToCS-based agent evaluation framework for ADF
まだ誰も着手していません。
評価
調査の方向性
Start with cto-executive-system/plans/tocs-terraphim-ai-evaluation-plan.md, cto-executive-system/plans/adf-architecture-improvements.md, and the referenced ToCS paper and repository; review dependency #689 before planning hook-based work. Done requires the design to be split into scoped sub-issues covering the baseline, probe collection, scoring, belief monitoring, NightwatchMonitor integration, and later KG enrichment.
索引モデルが issue の本文から書いたものです。
説明
Context
No existing ADF mechanism measures whether agents genuinely understand the codebases they operate on, or merely execute surface-level patterns. Theory of Code Space (ToCS, arXiv:2603.00601) provides a 4-dimension evaluation framework for exactly this.
Proposal
Implement ToCS-inspired evaluation to measure ADF agent effectiveness across four dimensions:
Evaluation Dimensions
| Dimension | What It Measures | Metric |
|---|---|---|
| Construct | Does the agent build an accurate dependency map? | Edge F1 by type (IMPORTS, CALLS_API, REGISTRY_WIRES, DATA_FLOWS_TO) |
| Revise | Does the agent update beliefs when code changes? | Belief revision score (delta accuracy after code change) |
| Exploit | Can the agent predict impact of changes? | Counterfactual probe accuracy |
| Constraints | Does the agent discover architectural rules? | Invariant discovery F1 vs CLAUDE.md/domain model rules |
Implementation Phases
- Phase 0: Run ToCS benchmark against terraphim-ai workspace with current agents (baseline)
- Phase 1: Add periodic cognitive map probing -- every N tool calls, externalise understanding as structured JSON
- Phase 2: Compare probes against ground truth (KG-derived dependency graph) to compute scores
- Phase 3: Feed scores to NightwatchMonitor as new signal type (alert on degradation)
Cognitive Map Probing
- Injected via
PreToolUsehooks (Agent SDK) or system messages (subprocess) - Agent outputs structured JSON: nodes (modules), edges (dependencies, typed), confidence scores
- Compared against ground truth from terraphim KG + tree-sitter analysis
Key Insight from ToCS Research
- Aho-Corasick automata cover ~67% of edges (IMPORTS level)
- CALLS_API (~17%) and DATA_FLOWS_TO (~7%) require semantic understanding
- Some models show "catastrophic belief collapse" -- losing knowledge between probes
- Evaluation framework should be built BEFORE KG enrichment (measure first, improve later)
Sub-issues (to be created during design phase)
- Run ToCS baseline against terraphim-ai workspace
- Implement cognitive map probe injection and collection
- Implement belief stability monitoring (successive probe comparison)
- Integrate evaluation scores with NightwatchMonitor
- KG enrichment with tree-sitter call graph (after baseline confirms gap)
References
- ToCS paper: https://arxiv.org/abs/2603.00601
- ToCS repo: https://github.com/che-shr-cat/tocs
- KB article:
cto-executive-system/knowledge/external/context-engineering/tocs-theory-of-code-space-benchmark.md - Expansion plan:
cto-executive-system/plans/tocs-terraphim-ai-evaluation-plan.md - ADF plan:
cto-executive-system/plans/adf-architecture-improvements.md(item 3.1) - Depends on: #689 (Agent SDK migration for hook-based probe injection)
- Related: #682 (Pi eval epic), #687 (steering queues)
- 主要言語
- Rust
- スター
- 64
- フォーク
- 5
- 平均マージ
- 1時間 17分
- マージ済み PR(30日)
- 2
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
terraphim/terraphim-ai のほかの issue
-
難易度 5/5 1週間以上 初心者へのやさしさ 25/100
terraphim/terraphim-ai#885 ·
-
難易度 4/5 3〜5日 初心者へのやさしさ 55/100
terraphim/terraphim-ai#871 ·
-
enhancement
難易度 5/5 1週間以上 初心者へのやさしさ 25/100
terraphim/terraphim-ai#810 · コメント 2 件 ·
-
enhancement
難易度 5/5 1週間以上 初心者へのやさしさ 35/100
terraphim/terraphim-ai#729 ·
-
enhancement
難易度 5/5 1週間以上 初心者へのやさしさ 35/100
terraphim/terraphim-ai#728 ·
terraphim/terraphim-ai の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
メンテナーはふだん 1 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 90/100
メンテナーはふだん 1 日以内に返信
-
agent:triaged bug bughunt pm:npm priority:p1
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
SocketDev/socket-patch#464 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
メンテナーはふだん 1 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
メンテナーはふだん 1 日以内に返信