Evaluation: Include judge cost tooltip
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 2/5
- Tempo stimato
- 1-3 ore
- Idoneità per principianti
- 78/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Specificata chiaramente
- Stato di attività
- Attiva
- Stack tecnologico
- typescript
- Ambito
- frontend
Direzione di ricerca
Inizia dalla pagina /evaluations e segui il tipo EvalCost e il rendering del tooltip di EvalRunCard. Aggiungi la voce relativa al costo del giudice accanto ai costi di risposta e di embedding quando job.cost.judge è presente, quindi verifica che il tooltip visualizzi tutti i costi disponibili e corrisponda a total_cost_usd.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Is your feature request related to a problem?
The cost tooltip on /evaluations only lists Response generation and omits the judge cost, leading to discrepancies in the cost breakdown compared to the total displayed. This causes confusion for users trying to understand the complete cost.
Describe the solution you'd like
- Update
EvalCostto includejudge?: EvalCostEntry. - Render a Judge scoring entry in the tooltip of
EvalRunCardwhenjob.cost.judgeis present. - Ensure all cost entries (response, judge, embedding) are shown in the tooltip to match
total_cost_usd.
Additional Context
Backend response
{
"id": 887,
"run_name": "assistant_v2_v1_0_ai_cohort_2_evals_demo_goldenqna_1788945623629",
"dataset_name": "ai_cohort_2_evals_demo_goldenqna",
"config_id": "dc576d3c-5e86-4eef-9b95-6d1e2194cce4",
"config_version": 1,
"dataset_id": 709,
"batch_job_id": 1805,
"embedding_batch_job_id": null,
"status": "completed",
"run_mode": "fast",
"object_store_url": null,
"score_trace_url": "s3://ai-platform-documents-staging/3ce7b9fe-2900-4f33-9a68-8162568a41be/evaluations/score/887/traces_887.json",
"total_items": 9,
"score": {
"overall": {
"verdict": "Needs Refinement",
"breakdown": [
{
"key": "ground_truth",
"name": "Adherence to Ground Truth",
"delta": -0.45,
"score": 3.44,
"weight": 0.71,
"verdict": "Needs Refinement"
},
{
"key": "prompt",
"name": "Adherence to Prompt",
"delta": 1.11,
"score": 5,
"weight": 0.29,
"verdict": "Good"
}
],
"ai_summary": "**Overall read:** The run is in generally good shape — most questions score 4–5 on ground truth and a clean 5 on prompt adherence, with no KB in play. The model answers are consistently substantive and well-structured; the main tension is between the model giving richer, modern-science answers and golden answers that expect specific, textbook-narrow responses.\n\n**Top 3 to check:**\n\n**Question 9** — Ground-truth score of 0: the golden answer expects a very specific socio-demographic list (sex, skin colour, caste, mother tongue, etc.) but the model answered from a biological/population-genetics frame; this looks like a golden-dataset framing issue more than a model failure, but needs a human call on which answer the use case actually wants.\n\n**Question 7** — Borderline ground-truth score (2): the model explicitly refuses to classify by skin colour and race, directly conflicting with the golden answer that includes skin colour; this is a values/alignment tension between the model's safety behaviour and the expected answer — worth deciding whether the golden answer or the model's stance is appropriate for this context.\n\n**Question 4** — Minor: the macrophage-as-viral-factory stage (a key step in the reference answer) is omitted; solid overall but worth a quick check if curriculum accuracy to the specific textbook is required, pointing at the model.\n\nThese are go-verify pointers — open each item, read the actual answer against the use case requirements, and decide based on what the deployment needs.",
"overall_score": 3.89
},
"summary_scores": [
{
"avg": 3.44,
"std": 1.42,
"name": "Adherence to Ground Truth",
"data_type": "NUMERIC",
"total_pairs": 9
},
{
"avg": 5,
"std": 0,
"name": "Adherence to Prompt",
"data_type": "NUMERIC",
"total_pairs": 9
}
]
},
"unscoreable": null,
"is_score_updated": true,
"is_judge_run": true,
"cost": {
"judge": {
"model": "gpt-5.6-luna",
"cost_usd": 0.002832,
"input_tokens": 16061,
"total_tokens": 18104,
"output_tokens": 2043
},
"response": {
"model": "gpt-5.6-luna",
"cost_usd": 0.001942,
"input_tokens": 276,
"total_tokens": 3466,
"output_tokens": 3190
},
"total_cost_usd": 0.004774
},
"error_message": null,
"organization_id": 1,
"project_id": 1,
"inserted_at": "2026-09-09T09:20:25.182635",
"updated_at": "2026-09-09T09:24:23.885105"
},
- Lingua principale
- TypeScript
- Stelle
- 1
- Fork
- 0
- Merge medio
- 19h 4m
- PR unite (30g)
- 6
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Nessuna guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di ProjectTech4DevAI/kaapi-frontend
-
enhancement
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
ProjectTech4DevAI/kaapi-frontend#267 ·
I maintainer di solito rispondono entro 1 giorno
-
enhancement
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
ProjectTech4DevAI/kaapi-frontend#266 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 90/100
ProjectTech4DevAI/kaapi-frontend#261 ·
I maintainer di solito rispondono entro 1 giorno
-
Assessment: Include input columns in resultsForse già presa @vprashrex l’ha presa 11 giorni fa. Apertabug enhancement
ProjectTech4DevAI/kaapi-frontend#286 · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
enhancement
Difficoltà 4/5 3-5 giorni Idoneità per principianti 55/100
ProjectTech4DevAI/kaapi-frontend#274 ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di ProjectTech4DevAI/kaapi-frontend
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
I maintainer di solito rispondono entro 1 giorno
-
external-issue to-triage
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 90/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
diegosouzapw/OmniRoute#15401 ·
I maintainer di solito rispondono entro 2 giorni
-
Sign the pledgeAperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 95/100
input-output-hk/devx-updates#163 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
code-yeongyu/oh-my-openagent#9454 ·
I maintainer di solito rispondono entro 1 giorno