Arena-Hard-V2 judge: an unrecognized/garbled verdict silently scores as a tie and is never flagged failed
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 75/100
- Issue type
- Bug
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- python
- Domain
- testing-qa
Research direction
The issue is in llm-as-a-coach/scripts/eval/eval_gpt4o_fuzzy.py, specifically the _parse_arena_verdict and _arena_score_round functions. Start by reading the judge_arena_hard_v2 function to understand the flow. Define a set of valid verdicts (A>>B, A>B, A=B, B>A, B>>A) and update the failed check to validate against this set, similar to the judge_wildbench pattern. Run the existing tests to ensure the fix works.
Written by the indexing model from the issue text.
Description
llm-as-a-coach/scripts/eval/eval_gpt4o_fuzzy.py:389-450, aggregated at 716-746
_ARENA_VERDICT_PATTERNS = [r"\[\[([AB<>=]+)\]\]", r"\[([AB<>=]+)\]"]
def _parse_arena_verdict(judgment_text):
if not judgment_text:
return None
upper = judgment_text.upper()
for pattern in _ARENA_VERDICT_PATTERNS:
matches = re.findall(pattern, upper)
matches = [m for m in matches if m]
if matches:
return matches[-1].strip("\n")
return None
def _arena_score_round(verdict, flipped):
if verdict is None:
return 0.5
verdict = verdict.replace(" ", "")
if not flipped:
if verdict in ("B>A", "B>>A"): return 1.0
if verdict == "A=B": return 0.5
if verdict in ("A>B", "A>>B"): return 0.0
else:
if verdict in ("A>B", "A>>B"): return 1.0
if verdict == "A=B": return 0.5
if verdict in ("B>A", "B>>A"): return 0.0
return 0.5 # <- unrecognized-but-non-None verdict falls through here
def judge_arena_hard_v2(oai_client, sample, idx):
...
win_rate = (s1 + s2) / 2.0
failed = v1 is None or v2 is None # <- only catches fully-empty output
return idx, win_rate, {"round1": output1, "round2": output2}, failed
_parse_arena_verdict's character class [AB<>=]+ accepts any run of those five characters inside brackets, not just the five literals the judge prompt actually asks for (A>>B, A>B, A=B, B>A, B>>A). A judge reply like [[AB]] or [[A>A]] — no operator, or a self-comparison — passes the regex, so v1/v2 come back non-None, so failed stays False. But the string matches none of the five branches in _arena_score_round, so it falls through to the final return 0.5, the same value the function returns for a genuine, well-formed tie (A=B).
What happens (reproduced on the real module)
_parse_arena_verdict's permissive character class [AB<>=]+ lets any run of A/B/</>/= characters count as a 'parsed' verdict (e.g. [[AB]], [[A>A]]). That makes v1/v2 non-None, so the round is excluded from failed = v1 is None or v2 is None (eval_gpt4o_fuzzy.py:449), even though _arena_score_round (lines 404-424) has no branch for that string and falls through to the same return 0.5 used for a genuine A=B tie.
As a result, a broken-but-bracketed judge call is indistinguishable, in both avg_win_rate and failed_indices, from a real tie. The sibling judge_wildbench (line 334) already closes this gap by checking the parsed value against a known-good vocabulary (choice.strip() not in WILDBENCH_REWARD_MAP) instead of only checking for None.
Suggested fix: define the five valid literals and set failed = v1 not in VALID_VERDICTS or v2 not in VALID_VERDICTS, mirroring the existing in-file pattern.
Happy to open the PR.
- Dominant language
- Python
- Stars
- 4.5k
- Forks
- 380
- Avg merge
- 9m
- Merged PRs (30d)
- 1
Getting set up
We have not checked this project's setup files yet. Start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/LMOps
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
-
Difficulty 3/5 1-2 days Newbie friendliness 48/100
-
Can the current black-box GAD addition be directly used in Qwen 3.5?May be free again @YTianZHU claimed this 129 days ago, and no pull request is open. Open
-
【GAD】使用Qwen3-VL进行GAD训练极慢May be free again @YTianZHU claimed this 193 days ago, and no pull request is open. Open
-
[GAD] Question regarding Reward Score fluctuations and downstream evaluationMay be free again @YTianZHU claimed this 193 days ago, and no pull request is open. Open
Similar issues
-
docs pydanty:is-working
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
pydantic/pydantic-ai#8863 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
run-llama/llama_index#23278 ·
Maintainers usually reply within 2 days
-
documentation from-review-extraction github-actions priority: low severity:nit
Difficulty 1/5 Under an hour Newbie friendliness 92/100
LearningCircuit/local-deep-research#6946 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
oracle/langchain-oracle#323 ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 88/100
tenstorrent/tt-metal#58057 · 1 comment ·
Maintainers usually reply within 1 day