no_rope_layer_interval is declared but never applied, and SmolLM3's own config sets it
メンテナーはふだん 1 日以内に返信
評価
この issue はまだ評価されていません。
説明
ModelArgs declares no_rope_layer_interval, and examples/models/smollm3/3b_config.json sets it to 4, but nothing on the examples/models/llama export path reads it. On main (e4576d0) the name appears exactly once across the 32 Python files of that path:
examples/models/llama/model_args.py:119: no_rope_layer_interval: Optional[int] = (
__post_init__ does not touch it, and Attention.forward applies rope unconditionally (attention.py:529, attention.py:563) with no layer-index test. The two places that do honour it are backends/mlx/llm/et_attention.py:111 and examples/qualcomm/oss_scripts/llama/model/static_llama.py:288.
So a model that alternates RoPE and NoPE layers exports without complaint and comes out with rope on every layer.
Reproducing
examples/models/smollm3/3b_config.json and examples/models/smollm3/convert_weights.py are both in the tree, so SmolLM3-3B looks ready to export. Converting its weights and exporting through export_llm with that params json (XNNPACK, 8da4w + 8-bit embedding, 1852.2 MB) gives a .pte that loads and runs, then does this:
prompt: <|im_start|>user\nWhat is the capital of France?<|im_end|>\n<|im_start|>assistant\n
output: Okay, so you want to know what what what what what what what what what what what
what what what what what what what what what what what what what what what what
prompt: <|im_start|>user\nWhat is 17 times 4?<|im_end|>\n<|im_start|>assistant\n
output: I think you are looking for an answer that can be given that can be be be be be
determ determ determ determ determ determ determ determ determ determ#ae determ#
It starts as English and collapses, which is what a positional-encoding mismatch looks like: the first few tokens are fine and the drift grows with position.
What would help
Either implement the interval on this path, or reject a params json that sets a field the chosen backend does not implement. The silent version is the expensive one — the export succeeds, the file is the right size, and the model is simply a different model.
Measured with the executorch 1.4.0 wheel; the code paths quoted above are byte-identical on main at e4576d0.
cc @larryliu0820 @mergennachin @cccclai @helunwencser @jackzhxng @digantdesai
- 主要言語
- Python
- スター
- 5.1k
- フォーク
- 1.2k
- 平均マージ
- 2日 9時間
- マージ済み PR(30日)
- 661
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
pytorch/executorch のほかの issue
-
enhancement triaged
難易度 2/5 半日 初心者へのやさしさ 68/100
pytorch/executorch#21640 ·
メンテナーはふだん 1 日以内に返信
-
Prompts of max_seq_len tokens fail to prefill: export_llm bounds the KV-cache token input at max_seq_len - 1 but publishes get_max_seq_len = max_seq_len対応中かも @alpharomercoma が 1 日前に担当しました。 オープン
難易度 3/5 1〜2日 初心者へのやさしさ 20/100
pytorch/executorch#23678 ·
メンテナーはふだん 1 日以内に返信
-
[v1.6.0] Release Schedule and Tracker対応中かも @JacobSzwejbka が 2 日前に担当しました。 オープン
pytorch/executorch#23608 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
QNN sharded LLM export fails: ResolveDebugHandle is not the last edge pass (dep_table entry overwritten by SplitGraph registration)対応中かも @psiddh が 2 日前に担当しました。 オープンbug module: llm module: qnn triaged
pytorch/executorch#23580 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
難易度 4/5 3〜5日 初心者へのやさしさ 48/100
pytorch/executorch#23553 ·
メンテナーはふだん 1 日以内に返信
pytorch/executorch の issue をすべて見る
似ている issue
-
namespace operations
難易度 1/5 1時間未満 初心者へのやさしさ 72/100
EclipseFdn/open-vsx.org#14043 ·
メンテナーはふだん 1 日以内に返信
-
feedback simulation workshop
難易度 2/5 1〜3時間 初心者へのやさしさ 73/100
githubnext/gh-aw-workshop#4455 ·
メンテナーはふだん 1 日以内に返信
-
Triage 🩺
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
メンテナーはふだん 1 日以内に返信
-
[BUG] Container scenario crashes without expected_recovery_time, kube DNS example uses retry_waitオープンneeds-triage
難易度 2/5 1〜3時間 初心者へのやさしさ 77/100
krkn-chaos/krkn#1627 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
NousResearch/hermes-agent#136483 ·
メンテナーはふだん 1 日以内に返信