Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[Bug Report] Qwen3.5 and Qwen3Next adapters don't set `rmsnorm_uses_offset`, so edited or backward-hooked norms drop the `1 +` in `(1 + weight)`

クローズ 初心者向け
#1,868 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
2/5
見積もり時間
1〜3時間
初心者へのやさしさ
62/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
python
領域
backend

調査の方向性

qwen3_5.py、qwen3_5_multimodal.py、qwen3_next.py のアダプター初期化処理を読み、issue に記載されているように、関連する設定フラグが super().__init__() の前に設定されている箇所を確認してください。まず、これらのアダプターが rmsnorm_uses_offset をどのように設定するかを追跡してください。issue によると、2つの MoE アダプターは一覧にあるアダプターを継承しています。forward の変更、backward hooks、LN-rule について、提案されている小規模モデルのテストを実行してください。これらのケースが成功し、hook を適用していない出力が引き続き Hugging Face と一致すれば完了です。

索引モデルが issue の本文から書いたものです。

説明

Describe the bug

In transformers, Qwen3_5RMSNorm, Qwen3_5MoeRMSNorm and Qwen3NextRMSNorm compute normalized * (1 + weight), with weight initialised to zeros. This is the Gemma convention. The adapters for these models never set cfg.rmsnorm_uses_offset. So whenever NormalizationBridge computes a norm's output itself, instead of calling the HF module, it multiplies by weight and drops the 1 +.

An unhooked forward pass and observe-only hooks, such as those run_with_cache adds, are correct, because they call the HF module. Three cases go wrong:

  1. A forward hook that edits hook_scale or hook_normalized. Since #1527, the native-autograd path rebuilds the output from the hooked values in _apply_weight_and_bias, which reads the flag. Even an identity edit, lambda t, hook: t.clone(), changes the logits.
  2. A backward hook on hook_scale or hook_normalized. The bridge then runs the norm through _python_norm_forward, so attaching a backward hook for gradient attribution changes the forward output.
  3. The LN-rule, enabled with use_relevance_rules(model, RelevanceRules(normalization=True)). _NativeLNRuleForward.backward scales the gradient by weight instead of 1 + weight. The forward output is unaffected.

Affected adapters:

  • Qwen3_5ArchitectureAdapter
  • Qwen3_5MoeArchitectureAdapter, which inherits from Qwen3_5ArchitectureAdapter
  • Qwen3NextArchitectureAdapter
  • Qwen3_5MultimodalArchitectureAdapter
  • Qwen3_5MoeMultimodalArchitectureAdapter, which inherits from Qwen3_5MultimodalArchitectureAdapter

Not affected:

  • Dense Qwen3, Qwen3-MoE and Qwen3-VL. Their norms compute weight * normalized, with weight initialised to ones.
  • Qwen3_5RMSNormGated, the GatedDeltaNet gated norm. It uses plain weight, and GatedDeltaNetBridge calls it directly.
  • The Qwen3.5 vision tower, whose norms the bridge does not wrap in NormalizationBridge.

Setting bridge.cfg.rmsnorm_uses_offset = True on a booted bridge only fixes ln_final. setup_blocks_bridge deep-copies the block template for each layer, config included, so each block's norms keep their own copy of the config.

Code example

The script below uses tiny random-weight models and runs on CPU, with no download. It moves the norm weights off their zero init, as a trained checkpoint would. It prints the relative logit error ‖hooked − clean‖ / ‖clean‖. The first three columns use an identity forward edit on each norm's hook_scale. The last column uses a backward hook on ln_final.hook_scale with no forward edit. Every column should be about 1e-7.

import warnings

import torch
import transformers
from tokenizers import Tokenizer, models, pre_tokenizers

from transformer_lens.model_bridge import TransformerBridge

warnings.filterwarnings("ignore")  # the bridge warns (by design) when it leaves HF's forward


def tiny_tokenizer():
    """In-memory word-level tokenizer, so nothing is downloaded."""
    vocab = {"<unk>": 0, "<pad>": 1, "<bos>": 2, "<eos>": 3, **{f"t{i}": i for i in range(4, 128)}}
    tok = Tokenizer(models.WordLevel(vocab=vocab, unk_token="<unk>"))
    tok.pre_tokenizer = pre_tokenizers.Whitespace()
    tokenizer = transformers.PreTrainedTokenizerFast(
        tokenizer_object=tok, unk_token="<unk>", pad_token="<pad>", bos_token="<bos>", eos_token="<eos>"
    )
    tokenizer.init_kwargs.update(name_or_path="tiny", add_bos_token=True)
    return tokenizer


COMMON = dict(
    vocab_size=128, hidden_size=64, intermediate_size=128, num_hidden_layers=2,
    num_attention_heads=4, num_key_value_heads=2, head_dim=16,
    linear_num_key_heads=2, linear_num_value_heads=4, linear_key_head_dim=16,
    linear_value_head_dim=16, linear_conv_kernel_dim=4,
    layer_types=["linear_attention", "full_attention"],
)
MOE = dict(num_experts=4, num_experts_per_tok=2, moe_intermediate_size=32,
           shared_expert_intermediate_size=32)
MODELS = {
    "Qwen3_5ForCausalLM": transformers.Qwen3_5TextConfig(**COMMON),
    "Qwen3_5MoeForCausalLM": transformers.Qwen3_5MoeTextConfig(**COMMON, **MOE),
    "Qwen3NextForCausalLM": transformers.Qwen3NextConfig(**COMMON, **MOE),
}
NORMS = ["blocks.0.ln1", "blocks.1.attn.q_norm", "ln_final"]


def rel_err(a, b):
    return ((a - b).norm() / b.norm()).item()


tokens = torch.randint(4, 128, (1, 8), generator=torch.Generator().manual_seed(0))
print(f"{'model':22} " + " ".join(f"{n:>21}" for n in NORMS) + f" {'ln_final bwd hook':>18}")
for arch, cfg in MODELS.items():
    cfg.architectures = [arch]
    torch.manual_seed(0)
    hf = getattr(transformers, arch)(cfg).eval()
    with torch.no_grad():  # trained checkpoints have norm weights near 0, not at their zero init
        for module in hf.modules():
            if type(module).__name__.endswith("RMSNorm"):
                module.weight.normal_(0, 0.3)
    bridge = TransformerBridge.boot_transformers(
        arch, hf_model=hf, tokenizer=tiny_tokenizer(), device="cpu", dtype=torch.float32
    )
    with torch.no_grad():
        clean = bridge(tokens)
        assert rel_err(clean, hf(tokens).logits) < 1e-5  # unhooked forward matches HF

        # A forward hook that returns a copy of hook_scale: changes nothing in value.
        errs = [
            rel_err(bridge.run_with_hooks(tokens, fwd_hooks=[(f"{n}.hook_scale", lambda t, hook: t.clone())]), clean)
            for n in NORMS
        ]
    # A backward hook only (no forward edit): the forward output should be unchanged.
    out = bridge.run_with_hooks(tokens, bwd_hooks=[("ln_final.hook_scale", lambda g, hook: g)])
    errs.append(rel_err(out.detach(), clean))
    print(f"{arch:22} " + " ".join(f"{e:21.1e}" for e in errs[:-1]) + f" {errs[-1]:18.1e}")

Output on dev at 6a9f67d9:

model                           blocks.0.ln1  blocks.1.attn.q_norm              ln_final  ln_final bwd hook
Qwen3_5ForCausalLM                   6.9e-02               2.4e-01               9.5e-01            9.5e-01
Qwen3_5MoeForCausalLM                1.0e-01               2.5e-01               9.5e-01            9.5e-01
Qwen3NextForCausalLM                 7.0e-02               2.4e-01               9.5e-01            9.5e-01

The unhooked forward matches HF to about 1e-7 for all three models, so the assert passes.

System Info

  • Installed from source: dev at 6a9f67d9, set up with uv sync
  • Linux, Python 3.12.14, torch 2.11.0+cu130 running on CPU, transformers 5.13.0
  • Also reproduced on the PyPI release 4.0.0 with transformers 5.17.0 and torch 2.14.0+cpu

Additional context

#1527 fixed #1526 by making edits on the native path rebuild the output with offset-aware weights. The Gemma adapters declare the offset, so they get the right scale. These adapters don't declare it.

The fix I propose sets the flag in each adapter's __init__ before super().__init__() builds the component mapping, next to the existing setattr(cfg, "gated_q_proj", True). That is one line each in qwen3_5.py, qwen3_5_multimodal.py and qwen3_next.py; the two MoE adapters inherit it. With that change, every column above drops to between 8.7e-08 and 1.8e-07.

An alternative is to set it once in Qwen3ArchitectureAdapter when hybrid=True. That is shorter, but it ties the norm convention to the layer layout, and the two only happen to line up today. I'm happy to do either.

I checked the other readers of the flag. The fold_ln paths in weight_processing.py, the batched-expert fold in transformer_bridge.py and the weight-processing benchmark all skip these adapters, because they set supports_fold_ln = False. On the tiny models, enable_compatibility_mode() gives the same logits with and without the flag, and log-probs match HF to 1e-6 either way. With the fix, the identity-final-norm check in direct_logit_attribution.py compares against zeros, which is correct for (1 + weight).

One separate question, which the fix doesn't touch. After enable_compatibility_mode(), ln_final.w and blocks.{i}.ln1.w on these models hold the raw offset, with values near 0, rather than the effective scale. The Gemma adapters declare ArithmeticTensorConversion(ADDITION, 1.0) for their norm weights, while the hybrid Qwen adapters declare no weight conversions. Should these models get the same conversion?

I have a branch with the fix and tests ready. The tests run tiny models through the edit, backward-hook and LN-rule cases, and they fail on dev and pass with the fix. I'm happy to open the PR.

Searches for rmsnorm_uses_offset, qwen3_5, Qwen3.5, Qwen3Next, hook_scale, RMSNorm, norm offset and 1 + weight found nothing similar. The closest are #1526/#1527 and #1652, a different Qwen3.5/Qwen3Next adapter bug.

I worked through this with Claude's help, and I've checked the results and run the repro myself.

Checklist
  • I have checked that there is no similar issue in the repo (required)
主要言語
Python
スター
3.9k
フォーク
708
平均マージ
1日 17時間
マージ済み PR(30日)
70

環境構築

Codespaces で開く

このプロジェクトの開発コンテナを、あなたの GitHub アカウントでブラウザ上に起動します。

  • Dockerfile・Docker Compose ファイルなし
  • プルリクエストのテンプレートあり
  • コントリビューションガイドなし

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

TransformerLensOrg/TransformerLens のほかの issue

TransformerLensOrg/TransformerLens の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。