Arena-Hard-V2 judge: an unrecognized/garbled verdict silently scores as a tie and is never flagged failed
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 2/5
- Tempo stimato
- 1-3 ore
- Idoneità per principianti
- 75/100
- Tipo di issue
- Bug
- Chiarezza
- Specificata chiaramente
- Stato di attività
- Attiva
- Stack tecnologico
- python
- Ambito
- testing-qa
Direzione di ricerca
Il problema si trova in llm-as-a-coach/scripts/eval/eval_gpt4o_fuzzy.py, specificamente nelle funzioni _parse_arena_verdict e _arena_score_round. Iniziate leggendo la funzione judge_arena_hard_v2 per comprenderne il flusso. Definite un insieme di verdetti validi (A>>B, A>B, A=B, B>A, B>>A) e aggiornate il controllo fallito per convalidarlo rispetto a questo insieme, simile al modello di judge_wildbench. Eseguite i test esistenti per assicurarvi che la correzione funzioni.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
llm-as-a-coach/scripts/eval/eval_gpt4o_fuzzy.py:389-450, aggregated at 716-746
_ARENA_VERDICT_PATTERNS = [r"\[\[([AB<>=]+)\]\]", r"\[([AB<>=]+)\]"]
def _parse_arena_verdict(judgment_text):
if not judgment_text:
return None
upper = judgment_text.upper()
for pattern in _ARENA_VERDICT_PATTERNS:
matches = re.findall(pattern, upper)
matches = [m for m in matches if m]
if matches:
return matches[-1].strip("\n")
return None
def _arena_score_round(verdict, flipped):
if verdict is None:
return 0.5
verdict = verdict.replace(" ", "")
if not flipped:
if verdict in ("B>A", "B>>A"): return 1.0
if verdict == "A=B": return 0.5
if verdict in ("A>B", "A>>B"): return 0.0
else:
if verdict in ("A>B", "A>>B"): return 1.0
if verdict == "A=B": return 0.5
if verdict in ("B>A", "B>>A"): return 0.0
return 0.5 # <- unrecognized-but-non-None verdict falls through here
def judge_arena_hard_v2(oai_client, sample, idx):
...
win_rate = (s1 + s2) / 2.0
failed = v1 is None or v2 is None # <- only catches fully-empty output
return idx, win_rate, {"round1": output1, "round2": output2}, failed
_parse_arena_verdict's character class [AB<>=]+ accepts any run of those five characters inside brackets, not just the five literals the judge prompt actually asks for (A>>B, A>B, A=B, B>A, B>>A). A judge reply like [[AB]] or [[A>A]] — no operator, or a self-comparison — passes the regex, so v1/v2 come back non-None, so failed stays False. But the string matches none of the five branches in _arena_score_round, so it falls through to the final return 0.5, the same value the function returns for a genuine, well-formed tie (A=B).
What happens (reproduced on the real module)
_parse_arena_verdict's permissive character class [AB<>=]+ lets any run of A/B/</>/= characters count as a 'parsed' verdict (e.g. [[AB]], [[A>A]]). That makes v1/v2 non-None, so the round is excluded from failed = v1 is None or v2 is None (eval_gpt4o_fuzzy.py:449), even though _arena_score_round (lines 404-424) has no branch for that string and falls through to the same return 0.5 used for a genuine A=B tie.
As a result, a broken-but-bracketed judge call is indistinguishable, in both avg_win_rate and failed_indices, from a real tie. The sibling judge_wildbench (line 334) already closes this gap by checking the parsed value against a known-good vocabulary (choice.strip() not in WILDBENCH_REWARD_MAP) instead of only checking for None.
Suggested fix: define the five valid literals and set failed = v1 not in VALID_VERDICTS or v2 not in VALID_VERDICTS, mirroring the existing in-file pattern.
Happy to open the PR.
- Lingua principale
- Python
- Stelle
- 4.5k
- Fork
- 380
- Merge medio
- 9m
- PR unite (30g)
- 1
Preparare l'ambiente
Non abbiamo ancora controllato i file di configurazione di questo progetto. Parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di microsoft/LMOps
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 35/100
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 48/100
-
Can the current black-box GAD addition be directly used in Qwen 3.5?Forse di nuovo libera @YTianZHU l’ha presa 130 giorni fa e non c’è nessuna pull request aperta. Aperta
-
【GAD】使用Qwen3-VL进行GAD训练极慢Forse di nuovo libera @YTianZHU l’ha presa 194 giorni fa e non c’è nessuna pull request aperta. Aperta
-
[GAD] Question regarding Reward Score fluctuations and downstream evaluationForse di nuovo libera @YTianZHU l’ha presa 194 giorni fa e non c’è nessuna pull request aperta. Aperta
Tutte le issue di microsoft/LMOps
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
spec-kitty/spec-kitty#5319 ·
I maintainer di solito rispondono entro 1 giorno
-
backend::vllm diffusion multimodal
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
openai/openai-agents-python#5229 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
I maintainer di solito rispondono entro 1 giorno