[Bug Report] Qwen3.5 and Qwen3Next adapters don't set `rmsnorm_uses_offset`, so edited or backward-hooked norms drop the `1 +` in `(1 + weight)`
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
調査の方向性
qwen3_5.py、qwen3_5_multimodal.py、qwen3_next.py のアダプター初期化処理を読み、issue に記載されているように、関連する設定フラグが super().__init__() の前に設定されている箇所を確認してください。まず、これらのアダプターが rmsnorm_uses_offset をどのように設定するかを追跡してください。issue によると、2つの MoE アダプターは一覧にあるアダプターを継承しています。forward の変更、backward hooks、LN-rule について、提案されている小規模モデルのテストを実行してください。これらのケースが成功し、hook を適用していない出力が引き続き Hugging Face と一致すれば完了です。
索引モデルが issue の本文から書いたものです。
説明
Describe the bug
In transformers, Qwen3_5RMSNorm, Qwen3_5MoeRMSNorm and Qwen3NextRMSNorm compute normalized * (1 + weight), with weight initialised to zeros. This is the Gemma convention. The adapters for these models never set cfg.rmsnorm_uses_offset. So whenever NormalizationBridge computes a norm's output itself, instead of calling the HF module, it multiplies by weight and drops the 1 +.
An unhooked forward pass and observe-only hooks, such as those run_with_cache adds, are correct, because they call the HF module. Three cases go wrong:
- A forward hook that edits
hook_scaleorhook_normalized. Since #1527, the native-autograd path rebuilds the output from the hooked values in_apply_weight_and_bias, which reads the flag. Even an identity edit,lambda t, hook: t.clone(), changes the logits. - A backward hook on
hook_scaleorhook_normalized. The bridge then runs the norm through_python_norm_forward, so attaching a backward hook for gradient attribution changes the forward output. - The LN-rule, enabled with
use_relevance_rules(model, RelevanceRules(normalization=True))._NativeLNRuleForward.backwardscales the gradient byweightinstead of1 + weight. The forward output is unaffected.
Affected adapters:
Qwen3_5ArchitectureAdapterQwen3_5MoeArchitectureAdapter, which inherits fromQwen3_5ArchitectureAdapterQwen3NextArchitectureAdapterQwen3_5MultimodalArchitectureAdapterQwen3_5MoeMultimodalArchitectureAdapter, which inherits fromQwen3_5MultimodalArchitectureAdapter
Not affected:
- Dense Qwen3, Qwen3-MoE and Qwen3-VL. Their norms compute
weight * normalized, with weight initialised to ones. Qwen3_5RMSNormGated, the GatedDeltaNet gated norm. It uses plainweight, andGatedDeltaNetBridgecalls it directly.- The Qwen3.5 vision tower, whose norms the bridge does not wrap in
NormalizationBridge.
Setting bridge.cfg.rmsnorm_uses_offset = True on a booted bridge only fixes ln_final. setup_blocks_bridge deep-copies the block template for each layer, config included, so each block's norms keep their own copy of the config.
Code example
The script below uses tiny random-weight models and runs on CPU, with no download. It moves the norm weights off their zero init, as a trained checkpoint would. It prints the relative logit error ‖hooked − clean‖ / ‖clean‖. The first three columns use an identity forward edit on each norm's hook_scale. The last column uses a backward hook on ln_final.hook_scale with no forward edit. Every column should be about 1e-7.
import warnings
import torch
import transformers
from tokenizers import Tokenizer, models, pre_tokenizers
from transformer_lens.model_bridge import TransformerBridge
warnings.filterwarnings("ignore") # the bridge warns (by design) when it leaves HF's forward
def tiny_tokenizer():
"""In-memory word-level tokenizer, so nothing is downloaded."""
vocab = {"<unk>": 0, "<pad>": 1, "<bos>": 2, "<eos>": 3, **{f"t{i}": i for i in range(4, 128)}}
tok = Tokenizer(models.WordLevel(vocab=vocab, unk_token="<unk>"))
tok.pre_tokenizer = pre_tokenizers.Whitespace()
tokenizer = transformers.PreTrainedTokenizerFast(
tokenizer_object=tok, unk_token="<unk>", pad_token="<pad>", bos_token="<bos>", eos_token="<eos>"
)
tokenizer.init_kwargs.update(name_or_path="tiny", add_bos_token=True)
return tokenizer
COMMON = dict(
vocab_size=128, hidden_size=64, intermediate_size=128, num_hidden_layers=2,
num_attention_heads=4, num_key_value_heads=2, head_dim=16,
linear_num_key_heads=2, linear_num_value_heads=4, linear_key_head_dim=16,
linear_value_head_dim=16, linear_conv_kernel_dim=4,
layer_types=["linear_attention", "full_attention"],
)
MOE = dict(num_experts=4, num_experts_per_tok=2, moe_intermediate_size=32,
shared_expert_intermediate_size=32)
MODELS = {
"Qwen3_5ForCausalLM": transformers.Qwen3_5TextConfig(**COMMON),
"Qwen3_5MoeForCausalLM": transformers.Qwen3_5MoeTextConfig(**COMMON, **MOE),
"Qwen3NextForCausalLM": transformers.Qwen3NextConfig(**COMMON, **MOE),
}
NORMS = ["blocks.0.ln1", "blocks.1.attn.q_norm", "ln_final"]
def rel_err(a, b):
return ((a - b).norm() / b.norm()).item()
tokens = torch.randint(4, 128, (1, 8), generator=torch.Generator().manual_seed(0))
print(f"{'model':22} " + " ".join(f"{n:>21}" for n in NORMS) + f" {'ln_final bwd hook':>18}")
for arch, cfg in MODELS.items():
cfg.architectures = [arch]
torch.manual_seed(0)
hf = getattr(transformers, arch)(cfg).eval()
with torch.no_grad(): # trained checkpoints have norm weights near 0, not at their zero init
for module in hf.modules():
if type(module).__name__.endswith("RMSNorm"):
module.weight.normal_(0, 0.3)
bridge = TransformerBridge.boot_transformers(
arch, hf_model=hf, tokenizer=tiny_tokenizer(), device="cpu", dtype=torch.float32
)
with torch.no_grad():
clean = bridge(tokens)
assert rel_err(clean, hf(tokens).logits) < 1e-5 # unhooked forward matches HF
# A forward hook that returns a copy of hook_scale: changes nothing in value.
errs = [
rel_err(bridge.run_with_hooks(tokens, fwd_hooks=[(f"{n}.hook_scale", lambda t, hook: t.clone())]), clean)
for n in NORMS
]
# A backward hook only (no forward edit): the forward output should be unchanged.
out = bridge.run_with_hooks(tokens, bwd_hooks=[("ln_final.hook_scale", lambda g, hook: g)])
errs.append(rel_err(out.detach(), clean))
print(f"{arch:22} " + " ".join(f"{e:21.1e}" for e in errs[:-1]) + f" {errs[-1]:18.1e}")
Output on dev at 6a9f67d9:
model blocks.0.ln1 blocks.1.attn.q_norm ln_final ln_final bwd hook
Qwen3_5ForCausalLM 6.9e-02 2.4e-01 9.5e-01 9.5e-01
Qwen3_5MoeForCausalLM 1.0e-01 2.5e-01 9.5e-01 9.5e-01
Qwen3NextForCausalLM 7.0e-02 2.4e-01 9.5e-01 9.5e-01
The unhooked forward matches HF to about 1e-7 for all three models, so the assert passes.
System Info
- Installed from source:
devat6a9f67d9, set up withuv sync - Linux, Python 3.12.14, torch 2.11.0+cu130 running on CPU, transformers 5.13.0
- Also reproduced on the PyPI release 4.0.0 with transformers 5.17.0 and torch 2.14.0+cpu
Additional context
#1527 fixed #1526 by making edits on the native path rebuild the output with offset-aware weights. The Gemma adapters declare the offset, so they get the right scale. These adapters don't declare it.
The fix I propose sets the flag in each adapter's __init__ before super().__init__() builds the component mapping, next to the existing setattr(cfg, "gated_q_proj", True). That is one line each in qwen3_5.py, qwen3_5_multimodal.py and qwen3_next.py; the two MoE adapters inherit it. With that change, every column above drops to between 8.7e-08 and 1.8e-07.
An alternative is to set it once in Qwen3ArchitectureAdapter when hybrid=True. That is shorter, but it ties the norm convention to the layer layout, and the two only happen to line up today. I'm happy to do either.
I checked the other readers of the flag. The fold_ln paths in weight_processing.py, the batched-expert fold in transformer_bridge.py and the weight-processing benchmark all skip these adapters, because they set supports_fold_ln = False. On the tiny models, enable_compatibility_mode() gives the same logits with and without the flag, and log-probs match HF to 1e-6 either way. With the fix, the identity-final-norm check in direct_logit_attribution.py compares against zeros, which is correct for (1 + weight).
One separate question, which the fix doesn't touch. After enable_compatibility_mode(), ln_final.w and blocks.{i}.ln1.w on these models hold the raw offset, with values near 0, rather than the effective scale. The Gemma adapters declare ArithmeticTensorConversion(ADDITION, 1.0) for their norm weights, while the hybrid Qwen adapters declare no weight conversions. Should these models get the same conversion?
I have a branch with the fix and tests ready. The tests run tiny models through the edit, backward-hook and LN-rule cases, and they fail on dev and pass with the fix. I'm happy to open the PR.
Searches for rmsnorm_uses_offset, qwen3_5, Qwen3.5, Qwen3Next, hook_scale, RMSNorm, norm offset and 1 + weight found nothing similar. The closest are #1526/#1527 and #1652, a different Qwen3.5/Qwen3Next adapter bug.
I worked through this with Claude's help, and I've checked the results and run the repro myself.
Checklist
- I have checked that there is no similar issue in the repo (required)
- 主要言語
- Python
- スター
- 3.9k
- フォーク
- 708
- 平均マージ
- 1日 17時間
- マージ済み PR(30日)
- 70
環境構築
このプロジェクトの開発コンテナを、あなたの GitHub アカウントでブラウザ上に起動します。
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドなし
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
TransformerLensOrg/TransformerLens のほかの issue
-
[Proposal] RoBERTa masked-LM adapter for TransformerBridge対応中かも @Canonik が今日担当しました。 オープンcomplexity-moderate new-architecture TransformerBridge
TransformerLensOrg/TransformerLens#1870 · コメント 1 件 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
[Proposal] Backward Lens: support gated MLP gate/ up/ down gradient factors対応中かも @janmenjayap が 7 日前に担当しました。 オープンcomplexity-moderate enhancement TransformerBridge
TransformerLensOrg/TransformerLens#1832 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
[Proposal] Sparse probing: optional groups argument so rows from one prompt can't straddle the split対応中かも @lorenzozanee が 13 日前に担当しました。 オープンcomplexity-simple enhancement help wanted TransformerBridge
難易度 4/5 3〜5日 初心者へのやさしさ 25/100
TransformerLensOrg/TransformerLens#1813 ·
メンテナーはふだん 1 日以内に返信
-
[Bug Report] _BLOCK_LIST_ATTRS hardcoded name list silently drops Raven's blocks from composition-score / head-label analysis対応中かも @LightWork666 が 17 日前に担当しました。 オープンbug complexity-moderate TransformerBridge
TransformerLensOrg/TransformerLens#1791 · コメント 2 件 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
[Proposal] SVD Circuits: singular-vector decomposition of a head's QK/ OV into causally-validated subfunctions対応中かも @janmenjayap が 28 日前に担当しました。 オープンcomplexity-high enhancement TransformerBridge
TransformerLensOrg/TransformerLens#1767 · コメント 3 件 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
TransformerLensOrg/TransformerLens の issue をすべて見る
似ている issue
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
Deepak3699/Ai_Mentor#244 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
btclib-org/btclib-node#1880 ·
メンテナーはふだん 1 日以内に返信
-
CONTRIBUTING.md: say how ticketless bug fixes and feature PRs are handled対応中かも @khuisman が今日担当しました。 オープンv0.9.2
難易度 1/5 1時間未満 初心者へのやさしさ 84/100
khuisman/mcp-gee-sweet#941 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
pyjanitor-devs/pyjanitor#1758 ·
メンテナーはふだん 1 日以内に返信
-
bug ready for review
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
odysseus-dev/odysseus#6641 ·
メンテナーはふだん 1 日以内に返信