Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Arena-Hard-V2 judge: an unrecognized/garbled verdict silently scores as a tie and is never flagged failed

Aperta Adatta ai principianti
#445 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
2/5
Tempo stimato
1-3 ore
Idoneità per principianti
75/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
python
Ambito
testing-qa

Direzione di ricerca

Il problema si trova in llm-as-a-coach/scripts/eval/eval_gpt4o_fuzzy.py, specificamente nelle funzioni _parse_arena_verdict e _arena_score_round. Iniziate leggendo la funzione judge_arena_hard_v2 per comprenderne il flusso. Definite un insieme di verdetti validi (A>>B, A>B, A=B, B>A, B>>A) e aggiornate il controllo fallito per convalidarlo rispetto a questo insieme, simile al modello di judge_wildbench. Eseguite i test esistenti per assicurarvi che la correzione funzioni.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

llm-as-a-coach/scripts/eval/eval_gpt4o_fuzzy.py:389-450, aggregated at 716-746

_ARENA_VERDICT_PATTERNS = [r"\[\[([AB<>=]+)\]\]", r"\[([AB<>=]+)\]"]

def _parse_arena_verdict(judgment_text):
    if not judgment_text:
        return None
    upper = judgment_text.upper()
    for pattern in _ARENA_VERDICT_PATTERNS:
        matches = re.findall(pattern, upper)
        matches = [m for m in matches if m]
        if matches:
            return matches[-1].strip("\n")
    return None

def _arena_score_round(verdict, flipped):
    if verdict is None:
        return 0.5
    verdict = verdict.replace(" ", "")
    if not flipped:
        if verdict in ("B>A", "B>>A"): return 1.0
        if verdict == "A=B": return 0.5
        if verdict in ("A>B", "A>>B"): return 0.0
    else:
        if verdict in ("A>B", "A>>B"): return 1.0
        if verdict == "A=B": return 0.5
        if verdict in ("B>A", "B>>A"): return 0.0
    return 0.5   # <- unrecognized-but-non-None verdict falls through here

def judge_arena_hard_v2(oai_client, sample, idx):
    ...
    win_rate = (s1 + s2) / 2.0
    failed = v1 is None or v2 is None   # <- only catches fully-empty output
    return idx, win_rate, {"round1": output1, "round2": output2}, failed

_parse_arena_verdict's character class [AB<>=]+ accepts any run of those five characters inside brackets, not just the five literals the judge prompt actually asks for (A>>B, A>B, A=B, B>A, B>>A). A judge reply like [[AB]] or [[A>A]] — no operator, or a self-comparison — passes the regex, so v1/v2 come back non-None, so failed stays False. But the string matches none of the five branches in _arena_score_round, so it falls through to the final return 0.5, the same value the function returns for a genuine, well-formed tie (A=B).

What happens (reproduced on the real module)

_parse_arena_verdict's permissive character class [AB<>=]+ lets any run of A/B/</>/= characters count as a 'parsed' verdict (e.g. [[AB]], [[A>A]]). That makes v1/v2 non-None, so the round is excluded from failed = v1 is None or v2 is None (eval_gpt4o_fuzzy.py:449), even though _arena_score_round (lines 404-424) has no branch for that string and falls through to the same return 0.5 used for a genuine A=B tie.

As a result, a broken-but-bracketed judge call is indistinguishable, in both avg_win_rate and failed_indices, from a real tie. The sibling judge_wildbench (line 334) already closes this gap by checking the parsed value against a known-good vocabulary (choice.strip() not in WILDBENCH_REWARD_MAP) instead of only checking for None.

Suggested fix: define the five valid literals and set failed = v1 not in VALID_VERDICTS or v2 not in VALID_VERDICTS, mirroring the existing in-file pattern.

Happy to open the PR.

Lingua principale
Python
Stelle
4.5k
Fork
380
Merge medio
9m
PR unite (30g)
1

Preparare l'ambiente

Non abbiamo ancora controllato i file di configurazione di questo progetto. Parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di microsoft/LMOps

Tutte le issue di microsoft/LMOps

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.