EIDOS episodic reasoning: prediction-outcome tracking for agent evolution
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 25/100
- Loại issue
- Tính năng
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Đình trệ
- Công nghệ
- rust
- Lĩnh vực
- ai, backend-api-design
Hướng nghiên cứu
Start by reading the terraphim_agent_evolution and terraphim_types crates, then review dependencies #597, #599, and #600. The affected areas also include terraphim_hooks and terraphim_mcp_server. Done means predictions and observed outcomes are modeled, stored and matched, accuracy is exposed, and confidence feedback is connected across the stated sources.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Summary
Implement episodic reasoning in terraphim_agent_evolution: every advisory/decision carries a predicted outcome. When the actual outcome is observed, prediction accuracy feeds back into future confidence scoring.
Motivation
Inspired by vibeship-spark-intelligence EIDOS (Episode-based Distillation) loop. terraphim_agent_evolution exists as a crate but is incomplete -- it's supposed to "track agent performance metrics, evolve agent capabilities over time." EIDOS provides the concrete mechanism.
EIDOS Loop
1. PREDICT: "This advisory/action will produce outcome X"
|
2. ACT: Agent or developer takes action
|
3. OBSERVE: Actual outcome Y recorded
|
4. EVALUATE: Compare prediction X vs outcome Y, compute accuracy
|
5. DISTILL: Update confidence for this class of advisory/action
|
(loop)
Concrete Application in Terraphim
Learning Capture Predictions
When terraphim-agent learn surfaces a past learning as advisory:
- Predict: "Following this learning will prevent error type Z"
- Observe: Did the developer follow the advice? Did the error recur?
- Evaluate: If followed and error didn't recur -> prediction validated
- Distill: Increase learning's reliability score
Judge Verdict Predictions
When the judge system issues a verdict:
- Predict: "This code change will cause issue type Z if merged"
- Observe: Was the verdict followed? Did the predicted issue manifest?
- Evaluate: If ignored and issue manifested -> judge was right but ignored
- Distill: Increase judge finding weight for this pattern
Hook Replacement Predictions
When terraphim_hooks replaces text (e.g., npm -> bun):
- Predict: "This replacement will succeed without side effects"
- Observe: Did the subsequent command succeed?
- Evaluate: If replacement caused a failure -> hook rule needs refinement
- Distill: Update hook confidence or add exception rule
Implementation Plan
1. Prediction type in terraphim_types
pub struct Prediction {
pub id: Ulid,
pub source: PredictionSource, // Learning, Judge, Hook, Agent
pub predicted_outcome: String,
pub confidence: f64, // 0.0 - 1.0
pub created_at: jiff::Timestamp,
pub observed_outcome: Option<ObservedOutcome>,
pub accuracy: Option<f64>,
}
2. Prediction tracking in terraphim_agent_evolution
- Store predictions in event store (#597)
- Match outcomes to predictions via correlation IDs
- Compute rolling accuracy per prediction source
- Expose accuracy metrics via MCP resource
3. Confidence feedback
- When prediction accuracy is high (>0.8 over 10+ predictions): increase source confidence
- When prediction accuracy is low (<0.5 over 10+ predictions): decrease source confidence or flag for review
- Confidence scores feed back into advisory ranking and judge dimensional scoring (#600)
Affected Crates
terraphim_agent_evolution(primary -- implement EIDOS loop)terraphim_types(add Prediction, ObservedOutcome types)terraphim_hooks(emit predictions for text replacements)terraphim_mcp_server(expose prediction accuracy metrics)
Dependencies
- #597 Event sourcing (predictions stored as events)
- #599 Enhanced learning capture (learning advisories generate predictions)
- #600 Dimensional verdict scoring (judge verdicts generate predictions)
Estimated Effort
~1 day for prediction types + basic tracking. Feedback loop refinement is ongoing.
Key Insight
Agent evolution is not about self-modifying code (Ouroboros pattern). It's about self-improving knowledge quality. The evolution happens in the advisory confidence scores, not in the source code. This is the right approach for terraphim's deterministic-first philosophy -- the Aho-Corasick engine stays constant, the confidence weights evolve.
- Ngôn ngữ chính
- Rust
- Star
- 62
- Fork
- 5
- Merge trung bình
- 2 giờ 27 phút
- Pull request đã merge (30 ngày)
- 1
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của terraphim/terraphim-ai
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 25/100
terraphim/terraphim-ai#885 ·
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 55/100
terraphim/terraphim-ai#871 ·
-
enhancement
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 25/100
terraphim/terraphim-ai#810 · 2 bình luận ·
-
enhancement
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
terraphim/terraphim-ai#729 ·
-
enhancement
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
terraphim/terraphim-ai#728 ·
Tất cả issue của terraphim/terraphim-ai
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
-
bug good first issue package: quic
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 78/100
-
`dora trace view` sends a non-canonical full UUID as-is, so a valid trace ID shows "No spans found" Đang mởcli coordinator rust
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
-
area: tasks enhancement good first issue help wanted
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
Jason-jo17/Polybench#15 · 1 bình luận ·