MiniCPM5-2B sampling defaults omit repeat-penalty; without it the model runs away on a large share of generations
还没有人认领这个 Issue。
评估
- 难度
- 3/5
- 预计耗时
- 1-2 天
- 新手友好度
- 58/100
- Issue 类型
- 文档
- 描述清晰度
- 基本清楚
- 活跃度
- 活跃
调研方向
首先,将 skills/minicpm5-deploy-llama-cpp/SKILL.md 和 docs/deployment/llama_cpp.md 中的采样表与已报告的 MiniCPM5-2B 设置和结果进行比较。在选择一个有文档记录的值之前,检查 repeat-penalty 是否会影响某项尚未测量的能力。完成的标准是两个指导位置保持一致,或者对省略有记录在案的理由。
由索引模型根据 Issue 内容生成。
描述
Summary
The sampling table in skills/minicpm5-deploy-llama-cpp/SKILL.md (and the matching guidance in docs/deployment/llama_cpp.md) lists only --temp 1.0 and --top-p 0.95 for MiniCPM5-2B Think. With exactly those settings and nothing else, I see a very high rate of generations where the thinking channel collapses into repetition and never terminates.
Adding a single flag, --repeat-penalty, changes the outcome dramatically. Everything else was held constant across all ten runs below; only that one value changed.
HumanEval+ (164 tasks), RTX 3060 12GB
| repeat-penalty | Q8_0 score | Q8_0 runaway rate | Q4_K_M score | Q4_K_M runaway rate |
|---|---|---|---|---|
| 1.00 (as documented) | 43.3 | 55.5% | 6.1 | 92.1% |
| 1.05 | 86.0 | 7.9% | 34.1 | 61.6% |
| 1.10 | 90.9 | 4.9% | 67.7 | 25.6% |
| 1.15 | 92.1 | 1.8% | 71.3 | 11.0% |
| 1.20 | 87.1 | 5.5% | 60.1 | 7.9% |
Both quants peak at 1.15 and regress at 1.20, so this is not simply "more is better" — 1.15 looks like a real optimum rather than an artifact.
Two things stand out:
- At the documented setting, Q4_K_M is effectively unusable for coding — 92% of tasks never produce an answer at all. The quant table in the same SKILL.md describes Q4_K_M as a "small drop, ideal for laptops", which is fair at 1.15 but very misleading at 1.00.
- The lower quant is far more sensitive to this. Q8_0 recovers almost fully in a single step from 1.00 to 1.05, while Q4_K_M needs the whole sweep. So the omission hurts exactly the users the recommended quant is aimed at.
Setup
- llama.cpp served through llama-swap (
ghcr.io/mostlygeek/llama-swap:unified-cuda), RTX 3060 12GB - Official GGUFs from
openbmb/MiniCPM5-2B-GGUF, both Q8_0 and Q4_K_M -ngl 99 -c 131072 -fa on --jinja -np 4 -ctk f16 -ctv f16 -b 2048 -ub 1024 --temp 1.0 --top-p 0.95- Thinking enabled. My harness runs one repair round, so these scores are not directly comparable to standard pass@1 numbers — the relative movement is the point, not the absolute values.
- "Runaway rate" is my harness killing a generation once it detects repetition. That is my own definition, not a MiniCPM concept.
Question
Was repeat-penalty left out deliberately — for example because it costs something on a capability I did not measure? If not, would you consider adding a recommended value to the sampling table for MiniCPM5-2B? I only tested HumanEval+ on one machine, so a value validated against your own eval suite would be much better than mine.
Related: #360 covered a different gap in the same sampling guidance, so it may be worth reviewing the documented profiles as a whole rather than patching one value at a time.
- 主要语言
- Jupyter Notebook
- 星标
- 11.1k
- 派生
- 766
- 平均合并
- 5 小时 3 分钟
- 30 天内合并 PR
- 6
环境准备
我们还没有检查这个项目的环境配置文件。先看它的 README,通用步骤见我们的新手贡献指南。
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
OpenBMB/MiniCPM 的其他 Issue
-
难度 1/5 1 小时以内 新手友好度 78/100
-
难度 5/5 一周以上 新手友好度 15/100
-
难度 3/5 1-2 天 新手友好度 55/100
-
feature
难度 4/5 3-5 天 新手友好度 35/100
-
难度 3/5 1-2 天 新手友好度 62/100
相似的 Issue
-
sync-en
难度 1/5 1-3 小时 新手友好度 88/100
维护者通常 2 天内回复
-
external
难度 2/5 1-3 小时 新手友好度 68/100
langchain-ai/docs#6255 ·
维护者通常 1 天内回复
-
detectors enhancement good first issue
难度 2/5 1-3 小时 新手友好度 86/100
SM260845/readme-gen#1 ·
-
难度 2/5 1-3 小时 新手友好度 88/100
angular/angularfire#3774 ·
维护者通常 2 天内回复
-
good first issue help wanted opensource september
难度 1/5 1 小时以内 新手友好度 88/100