Current bitnet-b1.58-2B-4T checkpoints contain all-zero MLP weight tensors (model generates garbage)
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 45/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Tranquilla
- Ambito
- machine-learning, tooling
Direzione di ricerca
Inizia eseguendo la riproduzione di safetensors con i file model.safetensors attuali e confrontali con i commit 9ff478e2487b e 9f43072f6949. Poi esamina utils/convert-ms-to-gguf-bitnet.py, in particolare SAFETENSORS_DATA_TYPES e le mappature ffn_sub_norm/attn_sub_norm. Il lavoro è completato quando è stato verificato che i checkpoint diversi da zero vengono ripristinati, che GGUF viene eseguito senza output spazzatura e che il percorso del convertitore interessato è utilizzabile.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Current bitnet-b1.58-2B-4T checkpoints contain all-zero MLP weight tensors (model generates garbage)
Summary
The current releases of microsoft/bitnet-b1.58-2B-4T, microsoft/bitnet-b1.58-2B-4T-bf16, and microsoft/bitnet-b1.58-2B-4T-gguf contain weight tensors that are entirely zero in several layers. Running the GGUF through bitnet.cpp produces repeated garbage tokens, consistent with the zeroed weights. The original releases (April 2025) are intact — the breakage was introduced by a later re-upload ("Update Model").
Affected repos / files (current main)
| Repo | File | Size |
|---|---|---|
microsoft/bitnet-b1.58-2B-4T |
model.safetensors (U8 "quantized" format) |
1,178,623,988 B |
microsoft/bitnet-b1.58-2B-4T-bf16 |
model.safetensors |
4,825,679,400 B |
microsoft/bitnet-b1.58-2B-4T-gguf |
ggml-model-i2_s.gguf |
1,187,801,280 B |
Evidence (verified 2026-08-10, direct tensor inspection via safetensors/torch)
microsoft/bitnet-b1.58-2B-4T-bf16 (current):
model.layers.0.mlp.down_proj.weight [2560, 6912] min=0 max=0 mean=0 nonzero=0
model.layers.0.mlp.gate_proj.weight [6912, 2560] min=0 max=0 mean=0 nonzero=0
model.layers.0.mlp.up_proj.weight [6912, 2560] min=0 max=0 mean=0 nonzero=0
model.layers.1.mlp.up_proj.weight nonzero=0 / 17,694,720
model.layers.10.mlp.up_proj.weight nonzero=0 / 17,694,720
model.layers.2.mlp.up_proj.weight nonzero=6,733,450 (has data — only some layers zeroed)
microsoft/bitnet-b1.58-2B-4T (current, U8):
model.layers.0.mlp.down_proj.weight uint8, all 0x00; weight_scale = 0.0
model.layers.0.mlp.gate_proj.weight uint8, all 0x00
model.layers.0.self_attn.q_proj.weight (valid 2-bit packed ternary data)
microsoft/bitnet-b1.58-2B-4T-gguf (current):
blk.0.ffn_up I2_S, nonzero=0 / 4,423,712
blk.2.ffn_up I2_S, nonzero=0 / 4,423,712
blk.1.ffn_up I2_S, nonzero=4,383,655 (has data)
Note the zeroed layers differ between the bf16 repo ({0,1,10}) and the GGUF repo ({0,2}) — the GGUF was evidently built from a different broken snapshot.
Reproduction (5 lines):
from safetensors import safe_open
with safe_open("model.safetensors", framework="pt") as f:
w = f.get_tensor("model.layers.0.mlp.up_proj.weight")
print((w != 0).sum().item(), "/", w.numel()) # -> 0 / 17694720
Running the current ggml-model-i2_s.gguf through bitnet.cpp (llama-cli, temp=0) outputs only ? tokens — the model is non-functional.
History / root cause
The original releases are intact and work:
microsoft/bitnet-b1.58-2B-4Tcommit9ff478e2487b→model.safetensors= 1,835,292,112 B (original)microsoft/bitnet-b1.58-2B-4T-ggufcommit9f43072f6949→ggml-model-i2_s.gguf= 1,844,472,032 B (original)
The re-upload commit 24edd43d41aa ("Update Model") replaced the file with the broken 1.18 GB version.
Secondary issue (tooling)
The new checkpoint format is not supported by any current converter:
- bitnet.cpp
utils/convert-ms-to-gguf-bitnet.pyfails withKeyError: 'U8'inSAFETENSORS_DATA_TYPES(newuint8dtype not registered) - upstream llama.cpp
convert_hf_to_gguf.pycannot map the newffn_sub_norm/attn_sub_normtensor names - Only the old-format files (original bf16 / original GGUF) can be converted and run today
Suggested fix
- Restore the correct weights (re-upload from
9ff478e2/9f43072f), or fix whatever produced the zeroed tensors - Regenerate the GGUF from a verified-good checkpoint
- Register the
U8dtype (andffn_sub_normmapping) in bitnet.cpp's converter so the new format is usable
- Lingua principale
- C++
- Stelle
- 40.3k
- Fork
- 3.7k
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di microsoft/BitNet
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
-
i2_s SIGSEGVs at n_ubatch >= 32: BLAS backend dequantises by a row stride 4x the real packed row Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 92/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
Tutte le issue di microsoft/BitNet
Issue simili
-
[CI] Nightly Clang build broken: structured binding captured in a lambda in FileDataSource.cpp Apertabug build
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 91/100
facebookincubator/velox#19194 ·
-
JIT-compiled number -> Decimal conversion silently overflows instead of raising DECIMAL_OVERFLOW Apertafuzz
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
ClickHouse/ClickHouse#122114 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
-
module/agent platform/macos type/bug/regression
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
-
enhancement PyCDE
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100