Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

EIDOS episodic reasoning: prediction-outcome tracking for agent evolution

Aperta
#601 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
5/5
Tempo stimato
Più di una settimana
Idoneità per principianti
25/100
Tipo di issue
Funzionalità
Chiarezza
Abbastanza chiara
Stato di attività
Ferma
Stack tecnologico
rust

Direzione di ricerca

Start by reading the terraphim_agent_evolution and terraphim_types crates, then review dependencies #597, #599, and #600. The affected areas also include terraphim_hooks and terraphim_mcp_server. Done means predictions and observed outcomes are modeled, stored and matched, accuracy is exposed, and confidence feedback is connected across the stated sources.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

architecture enhancement multi-agent

Summary

Implement episodic reasoning in terraphim_agent_evolution: every advisory/decision carries a predicted outcome. When the actual outcome is observed, prediction accuracy feeds back into future confidence scoring.

Motivation

Inspired by vibeship-spark-intelligence EIDOS (Episode-based Distillation) loop. terraphim_agent_evolution exists as a crate but is incomplete -- it's supposed to "track agent performance metrics, evolve agent capabilities over time." EIDOS provides the concrete mechanism.

EIDOS Loop

1. PREDICT: "This advisory/action will produce outcome X"
   |
2. ACT: Agent or developer takes action
   |
3. OBSERVE: Actual outcome Y recorded
   |
4. EVALUATE: Compare prediction X vs outcome Y, compute accuracy
   |
5. DISTILL: Update confidence for this class of advisory/action
   |
   (loop)

Concrete Application in Terraphim

Learning Capture Predictions

When terraphim-agent learn surfaces a past learning as advisory:

  • Predict: "Following this learning will prevent error type Z"
  • Observe: Did the developer follow the advice? Did the error recur?
  • Evaluate: If followed and error didn't recur -> prediction validated
  • Distill: Increase learning's reliability score
Judge Verdict Predictions

When the judge system issues a verdict:

  • Predict: "This code change will cause issue type Z if merged"
  • Observe: Was the verdict followed? Did the predicted issue manifest?
  • Evaluate: If ignored and issue manifested -> judge was right but ignored
  • Distill: Increase judge finding weight for this pattern
Hook Replacement Predictions

When terraphim_hooks replaces text (e.g., npm -> bun):

  • Predict: "This replacement will succeed without side effects"
  • Observe: Did the subsequent command succeed?
  • Evaluate: If replacement caused a failure -> hook rule needs refinement
  • Distill: Update hook confidence or add exception rule

Implementation Plan

1. Prediction type in terraphim_types
pub struct Prediction {
    pub id: Ulid,
    pub source: PredictionSource,     // Learning, Judge, Hook, Agent
    pub predicted_outcome: String,
    pub confidence: f64,              // 0.0 - 1.0
    pub created_at: jiff::Timestamp,
    pub observed_outcome: Option<ObservedOutcome>,
    pub accuracy: Option<f64>,
}
2. Prediction tracking in terraphim_agent_evolution
  • Store predictions in event store (#597)
  • Match outcomes to predictions via correlation IDs
  • Compute rolling accuracy per prediction source
  • Expose accuracy metrics via MCP resource
3. Confidence feedback
  • When prediction accuracy is high (>0.8 over 10+ predictions): increase source confidence
  • When prediction accuracy is low (<0.5 over 10+ predictions): decrease source confidence or flag for review
  • Confidence scores feed back into advisory ranking and judge dimensional scoring (#600)

Affected Crates

  • terraphim_agent_evolution (primary -- implement EIDOS loop)
  • terraphim_types (add Prediction, ObservedOutcome types)
  • terraphim_hooks (emit predictions for text replacements)
  • terraphim_mcp_server (expose prediction accuracy metrics)

Dependencies

  • #597 Event sourcing (predictions stored as events)
  • #599 Enhanced learning capture (learning advisories generate predictions)
  • #600 Dimensional verdict scoring (judge verdicts generate predictions)

Estimated Effort

~1 day for prediction types + basic tracking. Feedback loop refinement is ongoing.

Key Insight

Agent evolution is not about self-modifying code (Ouroboros pattern). It's about self-improving knowledge quality. The evolution happens in the advisory confidence scores, not in the source code. This is the right approach for terraphim's deterministic-first philosophy -- the Aho-Corasick engine stays constant, the confidence weights evolve.

Lingua principale
Rust
Stelle
62
Fork
5
Merge medio
2h 27m
PR unite (30g)
1

Guida per i contributori

Apri la guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di terraphim/terraphim-ai

Tutte le issue di terraphim/terraphim-ai

Issue simili

Altre issue su Rust

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.