Impediments to changing the dlrmv4 runs count
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 25/100
Research direction
Start by comparing the run-count and scoring rules in training_policies/training_rules.adoc with benchmark_meta.py and rcp_checker/rcp_checker.py. Then trace discard-count handling in rcp_checker.py and result_summarys.py, including the version and unet cases named in the issue. Done means the policy and code agree on the intended counts while preserving old-version results; the issue leaves some policy choices unresolved.
Written by the indexing model from the issue text.
Description
in training_policies/training_rules.adoc:
- run-count table (~line 565): change the minimum runs from 10 to 12
- scoring rule (line 570) says drop one fastest and one slowest. We should generalize to say "drop ceil(N/10) fastest and ceil(N/10) slowest". That will cover the old unet3d benchmark (which had N=40 and drop 4 highest and lowest), as well as all the N=5 and N=10 (where ceil(N/10) works out to drop 1 highest and lowest), and would cover dropping two highest and two lowest when N=12.
- The RCP rule requires references to provide 2N convergence numbers, so we either need to give an exception or we need to increase the number of RCP convergence numbers from 20 to 24.
logging
- The code does not follow DRY, so we need to modify both benchmark_meta.py
_ALL_RESULT_FILE_COUNTand rcp_checker/rcp_checker.pysubmission_runs(or better: have the rcp_checker import benchmark_meta.py). The contents seem to be identical, so this should be safe. _ALL_RESULT_FILE_COUNTandsubmission_runsare not versioned. So once a benchmark has a particular count, we can't change it in future versions without messing up how the checkers and result summarizers work on old versions.- the discard count is hardcoded to 1 all over the place with special exceptions for unet in some places (and bugs in other places where the unet exceptions are just ignored). Fixing this in such a way that old unet results don't change will require hardcoding a special exception for rounds prior to 6.1 at rcp_checker.py:406 (which doesn't handle unet correctly), but the problem at result_summarys.py:564 was introduced in April of 2026 and should be fixed (generalized) because it used to handle unet but then was hardcoded to 1 in April, so broke old unet results.
And additionally: there was a proposal to make the discard asymetric (discard two highest and only one lowest). And adding that to the code requires additional generalization in about a dozen places where it's currently hardcoded that the upper and lower discard count is the same.
- Dominant language
- Python
- Stars
- 43
- Forks
- 60
- Avg merge
- 2d 14h
- Merged PRs (30d)
- 1
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from mlcommons/logging
-
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
-
results summarizer is broken since 4.1.45Possibly taken @pgmpablo157321 claimed this 2 days ago. Open
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
-
Difficulty 4/5 3-5 days Newbie friendliness 25/100
All issues in mlcommons/logging
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
NousResearch/hermes-agent#136483 ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 88/100
Maintainers usually reply within 1 day
-
[BUG] LazyStackedTensorDictStore zeroes the last byte of a new key set on the last elementPossibly taken @peterdsharpe claimed this today. Openbug
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
pytorch/tensordict#2307 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Maintainers usually reply within 1 day
-
GrokModel.generate/a_generate pass an OpenAI-style list-of-dicts to xai_sdk.chat.user(), so every call crashes with a protobuf TypeError before any network I/OPossibly taken @Christian-Sidak claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
confident-ai/deepeval#3436 · 1 comment ·
Maintainers usually reply within 1 day