Current bitnet-b1.58-2B-4T checkpoints contain all-zero MLP weight tensors (model generates garbage)
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 45/100
- Issue 类型
- 缺陷
- 描述清晰度
- 基本清楚
- 活跃度
- 冷清
调研方向
首先,针对当前的 model.safetensors 文件运行 safetensors 复现,并将其与提交 9ff478e2487b 和 9f43072f6949 进行比较。然后检查 utils/convert-ms-to-gguf-bitnet.py,尤其是 SAFETENSORS_DATA_TYPES 以及 ffn_sub_norm/attn_sub_norm 映射。完成的标准是:已验证非零 checkpoint 得到恢复,GGUF 运行时不会产生垃圾输出,并且受影响的转换器路径可用。
由索引模型根据 Issue 内容生成。
描述
Current bitnet-b1.58-2B-4T checkpoints contain all-zero MLP weight tensors (model generates garbage)
Summary
The current releases of microsoft/bitnet-b1.58-2B-4T, microsoft/bitnet-b1.58-2B-4T-bf16, and microsoft/bitnet-b1.58-2B-4T-gguf contain weight tensors that are entirely zero in several layers. Running the GGUF through bitnet.cpp produces repeated garbage tokens, consistent with the zeroed weights. The original releases (April 2025) are intact — the breakage was introduced by a later re-upload ("Update Model").
Affected repos / files (current main)
| Repo | File | Size |
|---|---|---|
microsoft/bitnet-b1.58-2B-4T |
model.safetensors (U8 "quantized" format) |
1,178,623,988 B |
microsoft/bitnet-b1.58-2B-4T-bf16 |
model.safetensors |
4,825,679,400 B |
microsoft/bitnet-b1.58-2B-4T-gguf |
ggml-model-i2_s.gguf |
1,187,801,280 B |
Evidence (verified 2026-08-10, direct tensor inspection via safetensors/torch)
microsoft/bitnet-b1.58-2B-4T-bf16 (current):
model.layers.0.mlp.down_proj.weight [2560, 6912] min=0 max=0 mean=0 nonzero=0
model.layers.0.mlp.gate_proj.weight [6912, 2560] min=0 max=0 mean=0 nonzero=0
model.layers.0.mlp.up_proj.weight [6912, 2560] min=0 max=0 mean=0 nonzero=0
model.layers.1.mlp.up_proj.weight nonzero=0 / 17,694,720
model.layers.10.mlp.up_proj.weight nonzero=0 / 17,694,720
model.layers.2.mlp.up_proj.weight nonzero=6,733,450 (has data — only some layers zeroed)
microsoft/bitnet-b1.58-2B-4T (current, U8):
model.layers.0.mlp.down_proj.weight uint8, all 0x00; weight_scale = 0.0
model.layers.0.mlp.gate_proj.weight uint8, all 0x00
model.layers.0.self_attn.q_proj.weight (valid 2-bit packed ternary data)
microsoft/bitnet-b1.58-2B-4T-gguf (current):
blk.0.ffn_up I2_S, nonzero=0 / 4,423,712
blk.2.ffn_up I2_S, nonzero=0 / 4,423,712
blk.1.ffn_up I2_S, nonzero=4,383,655 (has data)
Note the zeroed layers differ between the bf16 repo ({0,1,10}) and the GGUF repo ({0,2}) — the GGUF was evidently built from a different broken snapshot.
Reproduction (5 lines):
from safetensors import safe_open
with safe_open("model.safetensors", framework="pt") as f:
w = f.get_tensor("model.layers.0.mlp.up_proj.weight")
print((w != 0).sum().item(), "/", w.numel()) # -> 0 / 17694720
Running the current ggml-model-i2_s.gguf through bitnet.cpp (llama-cli, temp=0) outputs only ? tokens — the model is non-functional.
History / root cause
The original releases are intact and work:
microsoft/bitnet-b1.58-2B-4Tcommit9ff478e2487b→model.safetensors= 1,835,292,112 B (original)microsoft/bitnet-b1.58-2B-4T-ggufcommit9f43072f6949→ggml-model-i2_s.gguf= 1,844,472,032 B (original)
The re-upload commit 24edd43d41aa ("Update Model") replaced the file with the broken 1.18 GB version.
Secondary issue (tooling)
The new checkpoint format is not supported by any current converter:
- bitnet.cpp
utils/convert-ms-to-gguf-bitnet.pyfails withKeyError: 'U8'inSAFETENSORS_DATA_TYPES(newuint8dtype not registered) - upstream llama.cpp
convert_hf_to_gguf.pycannot map the newffn_sub_norm/attn_sub_normtensor names - Only the old-format files (original bf16 / original GGUF) can be converted and run today
Suggested fix
- Restore the correct weights (re-upload from
9ff478e2/9f43072f), or fix whatever produced the zeroed tensors - Regenerate the GGUF from a verified-good checkpoint
- Register the
U8dtype (andffn_sub_normmapping) in bitnet.cpp's converter so the new format is usable
- 主要语言
- C++
- 星标
- 40.3k
- 派生
- 3.7k
- PR 合并指标
- 30 天内没有已合并 PR
贡献指南
这个仓库没有索引到贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
microsoft/BitNet 的其他 Issue
-
难度 2/5 1-3 小时 新手友好度 86/100
-
i2_s SIGSEGVs at n_ubatch >= 32: BLAS backend dequantises by a row stride 4x the real packed row 未关闭
难度 2/5 1-3 小时 新手友好度 76/100
-
难度 2/5 1-3 小时 新手友好度 86/100
-
难度 1/5 1 小时以内 新手友好度 92/100
-
难度 2/5 1-3 小时 新手友好度 76/100
相似的 Issue
-
bug build
难度 1/5 1 小时以内 新手友好度 91/100
facebookincubator/velox#19194 ·
-
JIT-compiled number -> Decimal conversion silently overflows instead of raising DECIMAL_OVERFLOW 未关闭fuzz
难度 2/5 1-3 小时 新手友好度 82/100
ClickHouse/ClickHouse#122114 ·
-
难度 2/5 1-3 小时 新手友好度 84/100
-
module/agent platform/macos type/bug/regression
难度 2/5 1-3 小时 新手友好度 88/100
-
enhancement PyCDE
难度 2/5 1-3 小时 新手友好度 78/100