[Bug]: BitNet FFN uses SILU instead of ReLU² — perplexity 99.8 vs 17.1 on every CPU backend
@praneshnikhar y travaille déjà.
Depuis le 9/8/2026.
Évaluation
- Difficulté
- 1/5
- Temps estimé
- Moins d'une heure
- Accessibilité débutants
- 30/100
- Type d'issue
- Bug
- Clarté
- Clairement spécifiée
- Activité
- À l'abandon
- Stack technique
- cpp
- Domaine
- backend, machine-learning
Piste de recherche
Commencez par 3rdparty/llama.cpp/src/models/bitnet.cpp et examinez l’appel d’activation de BitNet FFN, puis examinez ggml-org/llama.cpp#25885 et l’état du sous-module épinglé. C’est terminé lorsque la correction upstream est intégrée à la dépendance épinglée et que les résultats rapportés de llama-perplexity reviennent dans la plage de référence documentée.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Description
Model: microsoft/BitNet-b1.58-2B-4T (ggml-model-i2_s.gguf, ggml-model-tl1.gguf)
Where: 3rdparty/llama.cpp submodule → src/models/bitnet.cpp (the pinned
isHuangXin/llama.cpp@390c3077), also present in current ggml-org/llama.cpp master.
TL;DR
The BitNet FFN sub-layer applies SILU to the gate activation, but BitNet
b1.58 uses squared ReLU (relu(x)²). This is a one-line graph bug
(LLM_FFN_SILU should be LLM_FFN_RELU_SQR). It does not crash — the model
runs and emits plausible-looking text — but it is numerically mis-calibrated,
which shows up unmistakably in perplexity:
| Backend / kernel | Pure upstream (SILU) | With ReLU² fix | Official reference |
|---|---|---|---|
| x86 / I2_S (Intel) | 99.8178 ± 0.906 | 17.1086 ± 0.129 | 17.1090 ± 0.128 |
| x86 / I2_S (AMD) | 99.8178 ± 0.906 (bit-identical to Intel) | (identical by construction) | — |
| ARM / TL1 | 78.6877 ± 0.730 | 14.9641 ± 0.114 | (correct band) |
The Intel fixed value lands on the official 17.1090 to 4 decimal places
(Δ = 0.0004). The fix restores correctness on two independent kernel families
(scalar I2_S and LUT-based TL1) and across both x86 vendors, confirming this
is a graph-level bug, not a kernel/hardware issue.
Root cause
src/models/bitnet.cpp, in the per-layer FFN build:
cur = build_ffn(cur,
model.layers[il].ffn_up, NULL, model.layers[il].ffn_up_s,
model.layers[il].ffn_gate, NULL, model.layers[il].ffn_gate_s,
NULL, NULL, NULL,
NULL,
LLM_FFN_SILU, LLM_FFN_PAR, il); // ← BUG: BitNet b1.58 uses ReLU², not SILU
cb(cur, "ffn_sub_out", il);
The BitNet b1.58 architecture specifies squared ReLU in the FFN
(consistent with the reference model config). SILU here mis-scales activations
into the sub-norm, and the error compounds over 30 layers.
Fix (one line)
- LLM_FFN_SILU, LLM_FFN_PAR, il);
+ LLM_FFN_RELU_SQR, LLM_FFN_PAR, il);
LLM_FFN_RELU_SQR already exists in the enum (src/llama-graph.h), so this is
a pure activation-op swap with no new code. Throughput is unchanged (activation
op, not GEMM/LUT).
(For the origin fork Eddie-Wang1120/llama.cpp@bitnet, the same bug appears in
the pre-build_ffn monolithic form at llama.cpp:11905 as a literal
cur = ggml_silu(ctx0, cur); — replace with
cur = ggml_sqr(ctx0, ggml_relu(ctx0, cur));.)
Reproduction (clean-room, pure upstream)
Three fresh Azure VMs (Intel x86, AMD x86, ARM Ampere), pure upstream
microsoft/BitNet with recursive submodules, official 2B-4T model, standard
llama-perplexity over the same corpus:
git clone --recursive https://github.com/microsoft/BitNet.git && cd BitNet
python setup_env.py -md models/BitNet-b1.58-2B-4T -q i2_s # (or -q tl1 on ARM)
./build/bin/llama-perplexity -m models/BitNet-b1.58-2B-4T/ggml-model-i2_s.gguf \
-f <wikitext-2-raw/wiki.test.raw> -c 2048
# -> PPL ≈ 99.82 (x86 I2_S) / ≈ 78.69 (ARM TL1) [BROKEN]
Apply the one-line fix in 3rdparty/llama.cpp/src/models/bitnet.cpp, rebuild,
re-run:
# -> PPL 17.1086 (x86 I2_S) / 14.9641 (ARM TL1) [RESTORED — matches official 17.1090]
Both broken values are far outside the correct band; both fixes land inside it.
The broken PPL differs by kernel (99.82 vs 78.69) because the mis-calibrated
activation interacts with each kernel's numerics differently — but the
brokenness itself is universal and deterministic.
Why this hasn't been pinned before
This is distinct from the existing garbage-output reports:
- Not #547 / #305 / #470 / PR #580 (missing scalar I2_S fallback on
non-AVX2 x86 → empty kernel → constantGGGG; compile-time / hardware-gated). - Not #585 / PR #469 (wrong ARM 64/16 vs 128/32 weight-group layout →
scrambled weights →!!!!; ARM-CPU-specific).
Unlike those, this bug reproduces on every backend and both x86 vendors,
does not crash or emit a constant token, and only reveals itself as a
perplexity regression (~6× worse) with plausible-but-degraded text. That is
likely why the vague quality complaints — #216 ("local model doesn't have the
same quality as online"), #348, #106, #115 — were never root-caused: nobody
measured perplexity. This bug is a strong candidate explanation for those.
Suggested routing
The fix belongs in the pinned llama.cpp (isHuangXin/llama.cpp
release-bitnet-embedding-0.6b-270m, and canonically ggml-org/llama.cpp
master). Once merged upstream, microsoft/BitNet needs only a
3rdparty/llama.cpp submodule bump. Happy to open the upstream PR (patch
verified to apply clean at the pinned commit) and link it here.
Environment
- Model:
microsoft/BitNet-b1.58-2B-4Tofficial GGUF (I2_S, TL1) - llama.cpp submodule:
isHuangXin/llama.cpp@390c3077 - Build:
355c0c4d1 (9916), clang-18, cmake - VMs: Intel x86 (AVX-512), AMD x86 (AVX-512), ARM Ampere Altra — westus3
- Cross-vendor: Intel & AMD produced bit-identical broken PPL (99.8178)
Upstream fix: the canonical fix has been opened against ggml-org/llama.cpp master: ggml-org/llama.cpp#25885 (LLM_FFN_SILU → LLM_FFN_RELU_SQR in src/models/bitnet.cpp). Recommend re-syncing the pinned llama.cpp submodule once it lands.
- Langage dominant
- C++
- Étoiles
- 40.6k
- Forks
- 3.8k
- Métriques de merge des PR
- Aucune PR mergée en 30 j
Préparer son environnement
Ce projet ne fournit ni conteneur de développement, ni Dockerfile, ni guide de contribution : l'installation est à votre charge. Commencez par son README, et consultez notre guide de la première contribution pour les étapes générales.
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Autres issues de microsoft/BitNet
-
Difficulté 2/5 1-3 heures Accessibilité débutants 86/100
-
i2_s SIGSEGVs at n_ubatch >= 32: BLAS backend dequantises by a row stride 4x the real packed rowOuverte
Difficulté 2/5 1-3 heures Accessibilité débutants 76/100
-
Difficulté 2/5 1-3 heures Accessibilité débutants 86/100
-
Warning: `huggingface-cli` is deprecated and no longer works. Use `hf` instead.Peut-être pris @ousamabenyounes l’a pris il y a 61 jours. Ouverte
Difficulté 1/5 Moins d'une heure Accessibilité débutants 92/100
-
Difficulté 2/5 1-3 heures Accessibilité débutants 76/100
Toutes les issues de microsoft/BitNet
Issues similaires
-
Difficulté 2/5 1-3 heures Accessibilité débutants 64/100
utopia-rise/godot-jvm#1004 ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 2/5 1-3 heures Accessibilité débutants 70/100
Les mainteneurs répondent en général sous 3 jours
-
Difficulté 2/5 1-3 heures Accessibilité débutants 70/100
Les mainteneurs répondent en général sous 1 jour
-
new contributor
Difficulté 2/5 1-3 heures Accessibilité débutants 65/100
OpenMS/OpenMS#10512 · 1 commentaire ·
Les mainteneurs répondent en général sous 1 jour
-
Component: R
Difficulté 2/5 1-3 heures Accessibilité débutants 72/100
Les mainteneurs répondent en général sous 1 jour