EIDOS episodic reasoning: prediction-outcome tracking for agent evolution
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 25/100
- Tipo de issue
- Nueva funcionalidad
- Claridad
- Bastante claro
- Estado de actividad
- Estancado
- Stack tecnológico
- rust
- Área
- ai, backend-api-design
Línea de trabajo
Start by reading the terraphim_agent_evolution and terraphim_types crates, then review dependencies #597, #599, and #600. The affected areas also include terraphim_hooks and terraphim_mcp_server. Done means predictions and observed outcomes are modeled, stored and matched, accuracy is exposed, and confidence feedback is connected across the stated sources.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Summary
Implement episodic reasoning in terraphim_agent_evolution: every advisory/decision carries a predicted outcome. When the actual outcome is observed, prediction accuracy feeds back into future confidence scoring.
Motivation
Inspired by vibeship-spark-intelligence EIDOS (Episode-based Distillation) loop. terraphim_agent_evolution exists as a crate but is incomplete -- it's supposed to "track agent performance metrics, evolve agent capabilities over time." EIDOS provides the concrete mechanism.
EIDOS Loop
1. PREDICT: "This advisory/action will produce outcome X"
|
2. ACT: Agent or developer takes action
|
3. OBSERVE: Actual outcome Y recorded
|
4. EVALUATE: Compare prediction X vs outcome Y, compute accuracy
|
5. DISTILL: Update confidence for this class of advisory/action
|
(loop)
Concrete Application in Terraphim
Learning Capture Predictions
When terraphim-agent learn surfaces a past learning as advisory:
- Predict: "Following this learning will prevent error type Z"
- Observe: Did the developer follow the advice? Did the error recur?
- Evaluate: If followed and error didn't recur -> prediction validated
- Distill: Increase learning's reliability score
Judge Verdict Predictions
When the judge system issues a verdict:
- Predict: "This code change will cause issue type Z if merged"
- Observe: Was the verdict followed? Did the predicted issue manifest?
- Evaluate: If ignored and issue manifested -> judge was right but ignored
- Distill: Increase judge finding weight for this pattern
Hook Replacement Predictions
When terraphim_hooks replaces text (e.g., npm -> bun):
- Predict: "This replacement will succeed without side effects"
- Observe: Did the subsequent command succeed?
- Evaluate: If replacement caused a failure -> hook rule needs refinement
- Distill: Update hook confidence or add exception rule
Implementation Plan
1. Prediction type in terraphim_types
pub struct Prediction {
pub id: Ulid,
pub source: PredictionSource, // Learning, Judge, Hook, Agent
pub predicted_outcome: String,
pub confidence: f64, // 0.0 - 1.0
pub created_at: jiff::Timestamp,
pub observed_outcome: Option<ObservedOutcome>,
pub accuracy: Option<f64>,
}
2. Prediction tracking in terraphim_agent_evolution
- Store predictions in event store (#597)
- Match outcomes to predictions via correlation IDs
- Compute rolling accuracy per prediction source
- Expose accuracy metrics via MCP resource
3. Confidence feedback
- When prediction accuracy is high (>0.8 over 10+ predictions): increase source confidence
- When prediction accuracy is low (<0.5 over 10+ predictions): decrease source confidence or flag for review
- Confidence scores feed back into advisory ranking and judge dimensional scoring (#600)
Affected Crates
terraphim_agent_evolution(primary -- implement EIDOS loop)terraphim_types(add Prediction, ObservedOutcome types)terraphim_hooks(emit predictions for text replacements)terraphim_mcp_server(expose prediction accuracy metrics)
Dependencies
- #597 Event sourcing (predictions stored as events)
- #599 Enhanced learning capture (learning advisories generate predictions)
- #600 Dimensional verdict scoring (judge verdicts generate predictions)
Estimated Effort
~1 day for prediction types + basic tracking. Feedback loop refinement is ongoing.
Key Insight
Agent evolution is not about self-modifying code (Ouroboros pattern). It's about self-improving knowledge quality. The evolution happens in the advisory confidence scores, not in the source code. This is the right approach for terraphim's deterministic-first philosophy -- the Aho-Corasick engine stays constant, the confidence weights evolve.
- Lenguaje dominante
- Rust
- Estrellas
- 62
- Forks
- 5
- Merge medio
- 2 h 27 min
- PR fusionados (30 d)
- 1
Guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de terraphim/terraphim-ai
-
Dificultad 5/5 Más de una semana Aptitud para principiantes 25/100
terraphim/terraphim-ai#885 ·
-
Dificultad 4/5 3-5 días Aptitud para principiantes 55/100
terraphim/terraphim-ai#871 ·
-
enhancement
Dificultad 5/5 Más de una semana Aptitud para principiantes 25/100
terraphim/terraphim-ai#810 · 2 comentarios ·
-
enhancement
Dificultad 5/5 Más de una semana Aptitud para principiantes 35/100
terraphim/terraphim-ai#729 ·
-
enhancement
Dificultad 5/5 Más de una semana Aptitud para principiantes 35/100
terraphim/terraphim-ai#728 ·
Todos los issues de terraphim/terraphim-ai
Issues similares
-
Browser (wasm) relay client cannot connect to relays whose URL has a trailing-dot FQDN hostname Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
n0-computer/iroh#4550 ·
-
impl detach for native Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
paritytech/zombienet-sdk#591 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
farion1231/cc-switch#7638 · 1 comentario ·
-
onnx-ir re-exports ModelProto and GraphProto but not NodeProto, AttributeProto and AttributeType Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100