Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Current bitnet-b1.58-2B-4T checkpoints contain all-zero MLP weight tensors (model generates garbage)

オープン
#608 コメント 3 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
45/100
issue の種類
バグ
明瞭さ
おおむね明確
活発さ
静か
技術スタック
cpp, python, pytorch

調査の方向性

まず、現在の model.safetensors ファイルに対して safetensors の再現を実行し、コミット 9ff478e2487b および 9f43072f6949 と比較します。次に utils/convert-ms-to-gguf-bitnet.py を調査します。特に SAFETENSORS_DATA_TYPES と ffn_sub_norm/attn_sub_norm のマッピングを確認してください。ゼロでないチェックポイントが復元され、GGUF が不要な出力なしで実行され、影響を受けるコンバーターパスが利用可能であることを検証できれば完了です。

索引モデルが issue の本文から書いたものです。

説明

Current bitnet-b1.58-2B-4T checkpoints contain all-zero MLP weight tensors (model generates garbage)

Summary

The current releases of microsoft/bitnet-b1.58-2B-4T, microsoft/bitnet-b1.58-2B-4T-bf16, and microsoft/bitnet-b1.58-2B-4T-gguf contain weight tensors that are entirely zero in several layers. Running the GGUF through bitnet.cpp produces repeated garbage tokens, consistent with the zeroed weights. The original releases (April 2025) are intact — the breakage was introduced by a later re-upload ("Update Model").

Affected repos / files (current main)
Repo File Size
microsoft/bitnet-b1.58-2B-4T model.safetensors (U8 "quantized" format) 1,178,623,988 B
microsoft/bitnet-b1.58-2B-4T-bf16 model.safetensors 4,825,679,400 B
microsoft/bitnet-b1.58-2B-4T-gguf ggml-model-i2_s.gguf 1,187,801,280 B
Evidence (verified 2026-08-10, direct tensor inspection via safetensors/torch)

microsoft/bitnet-b1.58-2B-4T-bf16 (current):

model.layers.0.mlp.down_proj.weight   [2560, 6912]  min=0 max=0 mean=0 nonzero=0
model.layers.0.mlp.gate_proj.weight   [6912, 2560]  min=0 max=0 mean=0 nonzero=0
model.layers.0.mlp.up_proj.weight     [6912, 2560]  min=0 max=0 mean=0 nonzero=0
model.layers.1.mlp.up_proj.weight     nonzero=0 / 17,694,720
model.layers.10.mlp.up_proj.weight    nonzero=0 / 17,694,720
model.layers.2.mlp.up_proj.weight     nonzero=6,733,450 (has data — only some layers zeroed)

microsoft/bitnet-b1.58-2B-4T (current, U8):

model.layers.0.mlp.down_proj.weight   uint8, all 0x00; weight_scale = 0.0
model.layers.0.mlp.gate_proj.weight   uint8, all 0x00
model.layers.0.self_attn.q_proj.weight  (valid 2-bit packed ternary data)

microsoft/bitnet-b1.58-2B-4T-gguf (current):

blk.0.ffn_up   I2_S, nonzero=0 / 4,423,712
blk.2.ffn_up   I2_S, nonzero=0 / 4,423,712
blk.1.ffn_up   I2_S, nonzero=4,383,655 (has data)

Note the zeroed layers differ between the bf16 repo ({0,1,10}) and the GGUF repo ({0,2}) — the GGUF was evidently built from a different broken snapshot.

Reproduction (5 lines):

from safetensors import safe_open
with safe_open("model.safetensors", framework="pt") as f:
    w = f.get_tensor("model.layers.0.mlp.up_proj.weight")
    print((w != 0).sum().item(), "/", w.numel())   # -> 0 / 17694720

Running the current ggml-model-i2_s.gguf through bitnet.cpp (llama-cli, temp=0) outputs only ? tokens — the model is non-functional.

History / root cause

The original releases are intact and work:

  • microsoft/bitnet-b1.58-2B-4T commit 9ff478e2487b → model.safetensors = 1,835,292,112 B (original)
  • microsoft/bitnet-b1.58-2B-4T-gguf commit 9f43072f6949 → ggml-model-i2_s.gguf = 1,844,472,032 B (original)

The re-upload commit 24edd43d41aa ("Update Model") replaced the file with the broken 1.18 GB version.

Secondary issue (tooling)

The new checkpoint format is not supported by any current converter:

  • bitnet.cpp utils/convert-ms-to-gguf-bitnet.py fails with KeyError: 'U8' in SAFETENSORS_DATA_TYPES (new uint8 dtype not registered)
  • upstream llama.cpp convert_hf_to_gguf.py cannot map the new ffn_sub_norm / attn_sub_norm tensor names
  • Only the old-format files (original bf16 / original GGUF) can be converted and run today
Suggested fix
  1. Restore the correct weights (re-upload from 9ff478e2 / 9f43072f), or fix whatever produced the zeroed tensors
  2. Regenerate the GGUF from a verified-good checkpoint
  3. Register the U8 dtype (and ffn_sub_norm mapping) in bitnet.cpp's converter so the new format is usable
主要言語
C++
スター
40.3k
フォーク
3.7k
PR マージ指標
30日以内にマージされた PR はありません

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

microsoft/BitNet のほかの issue

microsoft/BitNet の issue をすべて見る

似ている issue

C++ の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。