Self-Driving Eval: Add UI for iteration
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 55/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Stack tecnologico
- typescript
Direzione di ricerca
Inizia tracciando il pattern esistente di miglioramento dei prompt nella console e nel relativo webhook receiver, quindi esamina POST /api/v2/evaluations/iterations e i tipi di risposta documentati. Il lavoro è completato quando gli utenti possono avviare un ciclo di iterazione, visualizzarne lo stato e la cronologia dei round, identificare il round migliore e comprendere ogni stop_reason terminale e il valore di Knowledge Base per ciascun round.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Is your feature request related to a problem?
The v2 self-driving eval currently lacks a UI for starting or monitoring the prompt iteration loop, making it an API-only feature. This limits user accessibility and understanding of the process.
Describe the solution you'd like
- Add a "Run iteration loop" entry point:
dataset_id,experiment_name,config_id,config_version,max_rounds - Handle the
202response withEvaluationIterationRunImmediatePublicdetails (iteration_run_id,status) - Implement a webhook receiver for the
EvaluationIterationReportPubliccallback, aligned with the prompt-improvement pattern - Display round-by-round history, the best round, and terminal
stop_reasonvalues (ceiling_reached,max_rounds_reached,round_failed)
Original issue
Is your feature request related to a problem?
v2 ships a self-driving eval → improve-prompt → eval loop (POST /api/v2/evaluations/iterations). There is no way to start one or watch it from the console, so the feature is API-only today.
Describe the solution you'd like
- Add a "Run iteration loop" entry point:
dataset_id,experiment_name,config_id,config_version,max_rounds - Handle the
202response and itsEvaluationIterationRunImmediatePublichandle (iteration_run_id,status) - Add a webhook receiver for the
EvaluationIterationReportPubliccallback, following the prompt-improvement pattern - Show round-by-round history, the best round, and the terminal
stop_reason(ceiling_reached,max_rounds_reached,round_failed)
Comments: Stop-score is the mean of Adherence to Ground Truth and Adherence to Prompt; Knowledge Base is recorded per round for visibility only — worth reflecting in the UI so the numbers aren't confusing.
Parent: #265
- Lingua principale
- TypeScript
- Stelle
- 1
- Fork
- 0
- Merge medio
- 10h 40m
- PR unite (30g)
- 4
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di ProjectTech4DevAI/kaapi-frontend
-
enhancement
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
-
enhancement
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
ProjectTech4DevAI/kaapi-frontend#266 · 1 commento ·
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 90/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
ProjectTech4DevAI/kaapi-frontend#258 · 3 commenti ·
-
enhancement
Difficoltà 4/5 3-5 giorni Idoneità per principianti 68/100
Tutte le issue di ProjectTech4DevAI/kaapi-frontend
Issue simili
-
Browser Waiting for: Product Owner
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
getsentry/sentry-javascript#24577 · 1 commento ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
agilepathway/label-checker#640 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
copse-dev/agent-pane#2953 ·
-
[aw] Upgrade available Apertaagentic-workflows
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 85/100
githubnext/rig#534 ·
-
automation missing-model model-sync provider:pioneer
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
anomalyco/models.dev#7701 ·