Current bitnet-b1.58-2B-4T checkpoints contain all-zero MLP weight tensors (model generates garbage)
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 45/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Ít trao đổi
- Lĩnh vực
- machine-learning, tooling
Hướng nghiên cứu
Bắt đầu bằng cách chạy bản tái hiện safetensors với các tệp model.safetensors hiện tại và so sánh chúng với các commit 9ff478e2487b và 9f43072f6949. Sau đó kiểm tra utils/convert-ms-to-gguf-bitnet.py, đặc biệt là SAFETENSORS_DATA_TYPES và các ánh xạ ffn_sub_norm/attn_sub_norm. Được coi là hoàn tất khi đã xác minh các checkpoint khác không được khôi phục, GGUF chạy mà không tạo ra đầu ra rác và đường dẫn converter bị ảnh hưởng có thể sử dụng được.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Current bitnet-b1.58-2B-4T checkpoints contain all-zero MLP weight tensors (model generates garbage)
Summary
The current releases of microsoft/bitnet-b1.58-2B-4T, microsoft/bitnet-b1.58-2B-4T-bf16, and microsoft/bitnet-b1.58-2B-4T-gguf contain weight tensors that are entirely zero in several layers. Running the GGUF through bitnet.cpp produces repeated garbage tokens, consistent with the zeroed weights. The original releases (April 2025) are intact — the breakage was introduced by a later re-upload ("Update Model").
Affected repos / files (current main)
| Repo | File | Size |
|---|---|---|
microsoft/bitnet-b1.58-2B-4T |
model.safetensors (U8 "quantized" format) |
1,178,623,988 B |
microsoft/bitnet-b1.58-2B-4T-bf16 |
model.safetensors |
4,825,679,400 B |
microsoft/bitnet-b1.58-2B-4T-gguf |
ggml-model-i2_s.gguf |
1,187,801,280 B |
Evidence (verified 2026-08-10, direct tensor inspection via safetensors/torch)
microsoft/bitnet-b1.58-2B-4T-bf16 (current):
model.layers.0.mlp.down_proj.weight [2560, 6912] min=0 max=0 mean=0 nonzero=0
model.layers.0.mlp.gate_proj.weight [6912, 2560] min=0 max=0 mean=0 nonzero=0
model.layers.0.mlp.up_proj.weight [6912, 2560] min=0 max=0 mean=0 nonzero=0
model.layers.1.mlp.up_proj.weight nonzero=0 / 17,694,720
model.layers.10.mlp.up_proj.weight nonzero=0 / 17,694,720
model.layers.2.mlp.up_proj.weight nonzero=6,733,450 (has data — only some layers zeroed)
microsoft/bitnet-b1.58-2B-4T (current, U8):
model.layers.0.mlp.down_proj.weight uint8, all 0x00; weight_scale = 0.0
model.layers.0.mlp.gate_proj.weight uint8, all 0x00
model.layers.0.self_attn.q_proj.weight (valid 2-bit packed ternary data)
microsoft/bitnet-b1.58-2B-4T-gguf (current):
blk.0.ffn_up I2_S, nonzero=0 / 4,423,712
blk.2.ffn_up I2_S, nonzero=0 / 4,423,712
blk.1.ffn_up I2_S, nonzero=4,383,655 (has data)
Note the zeroed layers differ between the bf16 repo ({0,1,10}) and the GGUF repo ({0,2}) — the GGUF was evidently built from a different broken snapshot.
Reproduction (5 lines):
from safetensors import safe_open
with safe_open("model.safetensors", framework="pt") as f:
w = f.get_tensor("model.layers.0.mlp.up_proj.weight")
print((w != 0).sum().item(), "/", w.numel()) # -> 0 / 17694720
Running the current ggml-model-i2_s.gguf through bitnet.cpp (llama-cli, temp=0) outputs only ? tokens — the model is non-functional.
History / root cause
The original releases are intact and work:
microsoft/bitnet-b1.58-2B-4Tcommit9ff478e2487b→model.safetensors= 1,835,292,112 B (original)microsoft/bitnet-b1.58-2B-4T-ggufcommit9f43072f6949→ggml-model-i2_s.gguf= 1,844,472,032 B (original)
The re-upload commit 24edd43d41aa ("Update Model") replaced the file with the broken 1.18 GB version.
Secondary issue (tooling)
The new checkpoint format is not supported by any current converter:
- bitnet.cpp
utils/convert-ms-to-gguf-bitnet.pyfails withKeyError: 'U8'inSAFETENSORS_DATA_TYPES(newuint8dtype not registered) - upstream llama.cpp
convert_hf_to_gguf.pycannot map the newffn_sub_norm/attn_sub_normtensor names - Only the old-format files (original bf16 / original GGUF) can be converted and run today
Suggested fix
- Restore the correct weights (re-upload from
9ff478e2/9f43072f), or fix whatever produced the zeroed tensors - Regenerate the GGUF from a verified-good checkpoint
- Register the
U8dtype (andffn_sub_normmapping) in bitnet.cpp's converter so the new format is usable
- Ngôn ngữ chính
- C++
- Star
- 40.3k
- Fork
- 3.7k
- Chỉ số merge pull request
- Không có pull request nào được merge trong 30 ngày
Chuẩn bị môi trường
Dự án này không cung cấp dev container, Dockerfile hay hướng dẫn đóng góp, nên bạn cần tự thiết lập môi trường: hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của microsoft/BitNet
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
-
i2_s SIGSEGVs at n_ubatch >= 32: BLAS backend dequantises by a row stride 4x the real packed rowĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
Tất cả issue của microsoft/BitNet
Issue tương tự
-
Feature
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 65/100
Narezzurri/OpenVPN-Config-Manager#95 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bot-found bug priority: P1
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
madenvel/KalinkaPlayer#251 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
sqlitebrowser/sqlitebrowser#4208 ·
-
ROSES ROSES - Student Review
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
Maintainer thường phản hồi trong vòng 1 ngày
-
area/ysql kind/bug priority/medium
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
yugabyte/yugabyte-db#34552 ·
Maintainer thường phản hồi trong vòng 1 ngày