Self-Driving Eval: Add UI for iteration
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 55/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- typescript
Research direction
Start by tracing the existing prompt-improvement pattern in the console and its webhook receiver, then inspect POST /api/v2/evaluations/iterations and the documented response types. Done means users can start an iteration loop, see its status and round history, identify the best round, and understand each terminal stop_reason and the per-round Knowledge Base value.
Written by the indexing model from the issue text.
Description
Is your feature request related to a problem?
The v2 self-driving eval currently lacks a UI for starting or monitoring the prompt iteration loop, making it an API-only feature. This limits user accessibility and understanding of the process.
Describe the solution you'd like
- Add a "Run iteration loop" entry point:
dataset_id,experiment_name,config_id,config_version,max_rounds - Handle the
202response withEvaluationIterationRunImmediatePublicdetails (iteration_run_id,status) - Implement a webhook receiver for the
EvaluationIterationReportPubliccallback, aligned with the prompt-improvement pattern - Display round-by-round history, the best round, and terminal
stop_reasonvalues (ceiling_reached,max_rounds_reached,round_failed)
Original issue
Is your feature request related to a problem?
v2 ships a self-driving eval → improve-prompt → eval loop (POST /api/v2/evaluations/iterations). There is no way to start one or watch it from the console, so the feature is API-only today.
Describe the solution you'd like
- Add a "Run iteration loop" entry point:
dataset_id,experiment_name,config_id,config_version,max_rounds - Handle the
202response and itsEvaluationIterationRunImmediatePublichandle (iteration_run_id,status) - Add a webhook receiver for the
EvaluationIterationReportPubliccallback, following the prompt-improvement pattern - Show round-by-round history, the best round, and the terminal
stop_reason(ceiling_reached,max_rounds_reached,round_failed)
Comments: Stop-score is the mean of Adherence to Ground Truth and Adherence to Prompt; Knowledge Base is recorded per round for visibility only — worth reflecting in the UI so the numbers aren't confusing.
Parent: #265
- Dominant language
- TypeScript
- Stars
- 1
- Forks
- 0
- Avg merge
- 10h 40m
- Merged PRs (30d)
- 4
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from ProjectTech4DevAI/kaapi-frontend
-
enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
-
enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
ProjectTech4DevAI/kaapi-frontend#266 · 1 comment ·
-
Difficulty 1/5 Under an hour Newbie friendliness 90/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
ProjectTech4DevAI/kaapi-frontend#258 · 3 comments ·
-
enhancement
Difficulty 4/5 3-5 days Newbie friendliness 68/100
All issues in ProjectTech4DevAI/kaapi-frontend
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
copse-dev/agent-pane#2953 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Eynzof/Hermes-CN-Desktop#610 ·
-
bug clawsweeper:linked-pr-open clawsweeper:needs-live-repro clawsweeper:no-new-fix-pr impact:message-loss issue-rating: 🐚 platinum hermit P2 regression
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
-
enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
-
calcite-components needs triage refactor
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
Esri/calcite-design-system#15203 ·