Ground truth errors bias weaker scores across all models
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 35/100
- Issue type
- Bug
- Clarity
- Needs clarification
- Activity status
- Active
- Tech stack
- python
- Domain
- data, machine-learning, testing-qa
Research direction
Start by locating the frozen test corpus and the evaluation or judging entry points, then inspect examples 1c73243f-9814-4252-a6b5-af03062f39c4 and f69fb188-7bbc-4b17-80b0-177dd2b6bf18. Compare their samples and ground truths with the reported outputs; completion should include a documented review method and an evidence-based ground-truth error rate.
Written by the indexing model from the issue text.
Description
While using this to enhance a mutimodal pipeline w/ Qwen3.8-27B, we've noticed that while the test corpus remains reproducible for benchmarking due to the frozen state of the test examples, it does not make it reliable. It should probably not be used as unreviewed training data as far as an actual pass/fail for vision models accuracy.
A simple spot-review of examples which our own harness plus additional context could not pass, on review, shows the defect is often with the sample and/or the ground truth being inaccurate.
Examples included:
1c73243f-9814-4252-a6b5-af03062f39c4
f69fb188-7bbc-4b17-80b0-177dd2b6bf18
First, and only two, that I've checked out of thousands, but you can see the problem with this method. This paints frontier models and the other models tested in a potentially weaker light as to their true abilities.
From the paper:
"The majority of reviewers agreed with the automated judge 97.1% of the time. All three reviewers unanimously considered the judge wrong in 5/477 = 1.0% of cases"
More helpful study information would be the expected error rate of the ground truths. I could not find that metric or that any review was done.
- Dominant language
- Python
- Stars
- 11
- Forks
- 1
- PR merge metrics
- No merged PRs in 30d
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Tencent/DiffSpot
All issues in Tencent/DiffSpot
Similar issues
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
use-agent-os/agent-os#3314 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
BasedHardware/omi#15662 · 1 comment ·
-
documentation help wanted
Difficulty 2/5 1-3 hours Newbie friendliness 90/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
AiursoftWeb/AnduinOS-2#19 ·