no_rope_layer_interval is declared but never applied, and SmolLM3's own config sets it
Maintainer thường phản hồi trong vòng 1 ngày
Đánh giá
Issue này chưa được đánh giá.
Mô tả
ModelArgs declares no_rope_layer_interval, and examples/models/smollm3/3b_config.json sets it to 4, but nothing on the examples/models/llama export path reads it. On main (e4576d0) the name appears exactly once across the 32 Python files of that path:
examples/models/llama/model_args.py:119: no_rope_layer_interval: Optional[int] = (
__post_init__ does not touch it, and Attention.forward applies rope unconditionally (attention.py:529, attention.py:563) with no layer-index test. The two places that do honour it are backends/mlx/llm/et_attention.py:111 and examples/qualcomm/oss_scripts/llama/model/static_llama.py:288.
So a model that alternates RoPE and NoPE layers exports without complaint and comes out with rope on every layer.
Reproducing
examples/models/smollm3/3b_config.json and examples/models/smollm3/convert_weights.py are both in the tree, so SmolLM3-3B looks ready to export. Converting its weights and exporting through export_llm with that params json (XNNPACK, 8da4w + 8-bit embedding, 1852.2 MB) gives a .pte that loads and runs, then does this:
prompt: <|im_start|>user\nWhat is the capital of France?<|im_end|>\n<|im_start|>assistant\n
output: Okay, so you want to know what what what what what what what what what what what
what what what what what what what what what what what what what what what what
prompt: <|im_start|>user\nWhat is 17 times 4?<|im_end|>\n<|im_start|>assistant\n
output: I think you are looking for an answer that can be given that can be be be be be
determ determ determ determ determ determ determ determ determ determ#ae determ#
It starts as English and collapses, which is what a positional-encoding mismatch looks like: the first few tokens are fine and the drift grows with position.
What would help
Either implement the interval on this path, or reject a params json that sets a field the chosen backend does not implement. The silent version is the expensive one — the export succeeds, the file is the right size, and the model is simply a different model.
Measured with the executorch 1.4.0 wheel; the code paths quoted above are byte-identical on main at e4576d0.
cc @larryliu0820 @mergennachin @cccclai @helunwencser @jackzhxng @digantdesai
- Ngôn ngữ chính
- Python
- Star
- 5.1k
- Fork
- 1.2k
- Merge trung bình
- 2 ngày 9 giờ
- Pull request đã merge (30 ngày)
- 661
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của pytorch/executorch
-
enhancement triaged
Độ khó 2/5 Nửa ngày Mức phù hợp với người mới 68/100
pytorch/executorch#21640 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Prompts of max_seq_len tokens fail to prefill: export_llm bounds the KV-cache token input at max_seq_len - 1 but publishes get_max_seq_len = max_seq_lenCó thể đã có người làm @alpharomercoma đã nhận 1 ngày trước. Đang mở
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 20/100
pytorch/executorch#23678 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[v1.6.0] Release Schedule and TrackerCó thể đã có người làm @JacobSzwejbka đã nhận 2 ngày trước. Đang mở
pytorch/executorch#23608 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
QNN sharded LLM export fails: ResolveDebugHandle is not the last edge pass (dep_table entry overwritten by SplitGraph registration)Có thể đã có người làm @psiddh đã nhận 2 ngày trước. Đang mởbug module: llm module: qnn triaged
pytorch/executorch#23580 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 48/100
pytorch/executorch#23553 ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của pytorch/executorch
Issue tương tự
-
[BUG] Container scenario crashes without expected_recovery_time, kube DNS example uses retry_waitĐang mởneeds-triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 77/100
krkn-chaos/krkn#1627 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
NousResearch/hermes-agent#136483 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
Maintainer thường phản hồi trong vòng 1 ngày
-
[BUG] LazyStackedTensorDictStore zeroes the last byte of a new key set on the last elementCó thể đã có người làm @peterdsharpe đã nhận hôm nay. Đang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
pytorch/tensordict#2307 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 1 ngày