Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Arena-Hard-V2 judge: an unrecognized/garbled verdict silently scores as a tie and is never flagged failed

Open Beginner friendly
#445 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
75/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python
Domain
testing-qa

Research direction

The issue is in llm-as-a-coach/scripts/eval/eval_gpt4o_fuzzy.py, specifically the _parse_arena_verdict and _arena_score_round functions. Start by reading the judge_arena_hard_v2 function to understand the flow. Define a set of valid verdicts (A>>B, A>B, A=B, B>A, B>>A) and update the failed check to validate against this set, similar to the judge_wildbench pattern. Run the existing tests to ensure the fix works.

Written by the indexing model from the issue text.

Description

llm-as-a-coach/scripts/eval/eval_gpt4o_fuzzy.py:389-450, aggregated at 716-746

_ARENA_VERDICT_PATTERNS = [r"\[\[([AB<>=]+)\]\]", r"\[([AB<>=]+)\]"]

def _parse_arena_verdict(judgment_text):
    if not judgment_text:
        return None
    upper = judgment_text.upper()
    for pattern in _ARENA_VERDICT_PATTERNS:
        matches = re.findall(pattern, upper)
        matches = [m for m in matches if m]
        if matches:
            return matches[-1].strip("\n")
    return None

def _arena_score_round(verdict, flipped):
    if verdict is None:
        return 0.5
    verdict = verdict.replace(" ", "")
    if not flipped:
        if verdict in ("B>A", "B>>A"): return 1.0
        if verdict == "A=B": return 0.5
        if verdict in ("A>B", "A>>B"): return 0.0
    else:
        if verdict in ("A>B", "A>>B"): return 1.0
        if verdict == "A=B": return 0.5
        if verdict in ("B>A", "B>>A"): return 0.0
    return 0.5   # <- unrecognized-but-non-None verdict falls through here

def judge_arena_hard_v2(oai_client, sample, idx):
    ...
    win_rate = (s1 + s2) / 2.0
    failed = v1 is None or v2 is None   # <- only catches fully-empty output
    return idx, win_rate, {"round1": output1, "round2": output2}, failed

_parse_arena_verdict's character class [AB<>=]+ accepts any run of those five characters inside brackets, not just the five literals the judge prompt actually asks for (A>>B, A>B, A=B, B>A, B>>A). A judge reply like [[AB]] or [[A>A]] — no operator, or a self-comparison — passes the regex, so v1/v2 come back non-None, so failed stays False. But the string matches none of the five branches in _arena_score_round, so it falls through to the final return 0.5, the same value the function returns for a genuine, well-formed tie (A=B).

What happens (reproduced on the real module)

_parse_arena_verdict's permissive character class [AB<>=]+ lets any run of A/B/</>/= characters count as a 'parsed' verdict (e.g. [[AB]], [[A>A]]). That makes v1/v2 non-None, so the round is excluded from failed = v1 is None or v2 is None (eval_gpt4o_fuzzy.py:449), even though _arena_score_round (lines 404-424) has no branch for that string and falls through to the same return 0.5 used for a genuine A=B tie.

As a result, a broken-but-bracketed judge call is indistinguishable, in both avg_win_rate and failed_indices, from a real tie. The sibling judge_wildbench (line 334) already closes this gap by checking the parsed value against a known-good vocabulary (choice.strip() not in WILDBENCH_REWARD_MAP) instead of only checking for None.

Suggested fix: define the five valid literals and set failed = v1 not in VALID_VERDICTS or v2 not in VALID_VERDICTS, mirroring the existing in-file pattern.

Happy to open the PR.

Dominant language
Python
Stars
4.5k
Forks
380
Avg merge
9m
Merged PRs (30d)
1

Getting set up

We have not checked this project's setup files yet. Start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from microsoft/LMOps

All issues in microsoft/LMOps

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.