Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

EIDOS episodic reasoning: prediction-outcome tracking for agent evolution

Đang mở
#601 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
25/100
Loại issue
Tính năng
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Đình trệ
Công nghệ
rust
Lĩnh vực
ai, backend-api-design

Hướng nghiên cứu

Start by reading the terraphim_agent_evolution and terraphim_types crates, then review dependencies #597, #599, and #600. The affected areas also include terraphim_hooks and terraphim_mcp_server. Done means predictions and observed outcomes are modeled, stored and matched, accuracy is exposed, and confidence feedback is connected across the stated sources.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

architecture enhancement multi-agent

Summary

Implement episodic reasoning in terraphim_agent_evolution: every advisory/decision carries a predicted outcome. When the actual outcome is observed, prediction accuracy feeds back into future confidence scoring.

Motivation

Inspired by vibeship-spark-intelligence EIDOS (Episode-based Distillation) loop. terraphim_agent_evolution exists as a crate but is incomplete -- it's supposed to "track agent performance metrics, evolve agent capabilities over time." EIDOS provides the concrete mechanism.

EIDOS Loop

1. PREDICT: "This advisory/action will produce outcome X"
   |
2. ACT: Agent or developer takes action
   |
3. OBSERVE: Actual outcome Y recorded
   |
4. EVALUATE: Compare prediction X vs outcome Y, compute accuracy
   |
5. DISTILL: Update confidence for this class of advisory/action
   |
   (loop)

Concrete Application in Terraphim

Learning Capture Predictions

When terraphim-agent learn surfaces a past learning as advisory:

  • Predict: "Following this learning will prevent error type Z"
  • Observe: Did the developer follow the advice? Did the error recur?
  • Evaluate: If followed and error didn't recur -> prediction validated
  • Distill: Increase learning's reliability score
Judge Verdict Predictions

When the judge system issues a verdict:

  • Predict: "This code change will cause issue type Z if merged"
  • Observe: Was the verdict followed? Did the predicted issue manifest?
  • Evaluate: If ignored and issue manifested -> judge was right but ignored
  • Distill: Increase judge finding weight for this pattern
Hook Replacement Predictions

When terraphim_hooks replaces text (e.g., npm -> bun):

  • Predict: "This replacement will succeed without side effects"
  • Observe: Did the subsequent command succeed?
  • Evaluate: If replacement caused a failure -> hook rule needs refinement
  • Distill: Update hook confidence or add exception rule

Implementation Plan

1. Prediction type in terraphim_types
pub struct Prediction {
    pub id: Ulid,
    pub source: PredictionSource,     // Learning, Judge, Hook, Agent
    pub predicted_outcome: String,
    pub confidence: f64,              // 0.0 - 1.0
    pub created_at: jiff::Timestamp,
    pub observed_outcome: Option<ObservedOutcome>,
    pub accuracy: Option<f64>,
}
2. Prediction tracking in terraphim_agent_evolution
  • Store predictions in event store (#597)
  • Match outcomes to predictions via correlation IDs
  • Compute rolling accuracy per prediction source
  • Expose accuracy metrics via MCP resource
3. Confidence feedback
  • When prediction accuracy is high (>0.8 over 10+ predictions): increase source confidence
  • When prediction accuracy is low (<0.5 over 10+ predictions): decrease source confidence or flag for review
  • Confidence scores feed back into advisory ranking and judge dimensional scoring (#600)

Affected Crates

  • terraphim_agent_evolution (primary -- implement EIDOS loop)
  • terraphim_types (add Prediction, ObservedOutcome types)
  • terraphim_hooks (emit predictions for text replacements)
  • terraphim_mcp_server (expose prediction accuracy metrics)

Dependencies

  • #597 Event sourcing (predictions stored as events)
  • #599 Enhanced learning capture (learning advisories generate predictions)
  • #600 Dimensional verdict scoring (judge verdicts generate predictions)

Estimated Effort

~1 day for prediction types + basic tracking. Feedback loop refinement is ongoing.

Key Insight

Agent evolution is not about self-modifying code (Ouroboros pattern). It's about self-improving knowledge quality. The evolution happens in the advisory confidence scores, not in the source code. This is the right approach for terraphim's deterministic-first philosophy -- the Aho-Corasick engine stays constant, the confidence weights evolve.

Ngôn ngữ chính
Rust
Star
62
Fork
5
Merge trung bình
2 giờ 27 phút
Pull request đã merge (30 ngày)
1

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của terraphim/terraphim-ai

Tất cả issue của terraphim/terraphim-ai

Issue tương tự

Thêm issue về Rust

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.