Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Epic: ToCS-based agent evaluation framework for ADF

オープン
#691 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
25/100
issue の種類
機能追加
明瞭さ
説明が足りない
活発さ
停滞
技術スタック
rust
領域
ai, testing

調査の方向性

Start with cto-executive-system/plans/tocs-terraphim-ai-evaluation-plan.md, cto-executive-system/plans/adf-architecture-improvements.md, and the referenced ToCS paper and repository; review dependency #689 before planning hook-based work. Done requires the design to be split into scoped sub-issues covering the baseline, probe collection, scoring, belief monitoring, NightwatchMonitor integration, and later KG enrichment.

索引モデルが issue の本文から書いたものです。

説明

enhancement

Context

No existing ADF mechanism measures whether agents genuinely understand the codebases they operate on, or merely execute surface-level patterns. Theory of Code Space (ToCS, arXiv:2603.00601) provides a 4-dimension evaluation framework for exactly this.

Proposal

Implement ToCS-inspired evaluation to measure ADF agent effectiveness across four dimensions:

Evaluation Dimensions
Dimension What It Measures Metric
Construct Does the agent build an accurate dependency map? Edge F1 by type (IMPORTS, CALLS_API, REGISTRY_WIRES, DATA_FLOWS_TO)
Revise Does the agent update beliefs when code changes? Belief revision score (delta accuracy after code change)
Exploit Can the agent predict impact of changes? Counterfactual probe accuracy
Constraints Does the agent discover architectural rules? Invariant discovery F1 vs CLAUDE.md/domain model rules
Implementation Phases
  1. Phase 0: Run ToCS benchmark against terraphim-ai workspace with current agents (baseline)
  2. Phase 1: Add periodic cognitive map probing -- every N tool calls, externalise understanding as structured JSON
  3. Phase 2: Compare probes against ground truth (KG-derived dependency graph) to compute scores
  4. Phase 3: Feed scores to NightwatchMonitor as new signal type (alert on degradation)
Cognitive Map Probing
  • Injected via PreToolUse hooks (Agent SDK) or system messages (subprocess)
  • Agent outputs structured JSON: nodes (modules), edges (dependencies, typed), confidence scores
  • Compared against ground truth from terraphim KG + tree-sitter analysis
Key Insight from ToCS Research
  • Aho-Corasick automata cover ~67% of edges (IMPORTS level)
  • CALLS_API (~17%) and DATA_FLOWS_TO (~7%) require semantic understanding
  • Some models show "catastrophic belief collapse" -- losing knowledge between probes
  • Evaluation framework should be built BEFORE KG enrichment (measure first, improve later)
Sub-issues (to be created during design phase)
  • Run ToCS baseline against terraphim-ai workspace
  • Implement cognitive map probe injection and collection
  • Implement belief stability monitoring (successive probe comparison)
  • Integrate evaluation scores with NightwatchMonitor
  • KG enrichment with tree-sitter call graph (after baseline confirms gap)

References

  • ToCS paper: https://arxiv.org/abs/2603.00601
  • ToCS repo: https://github.com/che-shr-cat/tocs
  • KB article: cto-executive-system/knowledge/external/context-engineering/tocs-theory-of-code-space-benchmark.md
  • Expansion plan: cto-executive-system/plans/tocs-terraphim-ai-evaluation-plan.md
  • ADF plan: cto-executive-system/plans/adf-architecture-improvements.md (item 3.1)
  • Depends on: #689 (Agent SDK migration for hook-based probe injection)
  • Related: #682 (Pi eval epic), #687 (steering queues)
主要言語
Rust
スター
64
フォーク
5
平均マージ
1時間 17分
マージ済み PR(30日)
2

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

terraphim/terraphim-ai のほかの issue

terraphim/terraphim-ai の issue をすべて見る

似ている issue

Rust の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。