no_rope_layer_interval is declared but never applied, and SmolLM3's own config sets it
Los mantenedores suelen responder en 1 día
Evaluación
Este issue todavía no se ha evaluado.
Descripción
ModelArgs declares no_rope_layer_interval, and examples/models/smollm3/3b_config.json sets it to 4, but nothing on the examples/models/llama export path reads it. On main (e4576d0) the name appears exactly once across the 32 Python files of that path:
examples/models/llama/model_args.py:119: no_rope_layer_interval: Optional[int] = (
__post_init__ does not touch it, and Attention.forward applies rope unconditionally (attention.py:529, attention.py:563) with no layer-index test. The two places that do honour it are backends/mlx/llm/et_attention.py:111 and examples/qualcomm/oss_scripts/llama/model/static_llama.py:288.
So a model that alternates RoPE and NoPE layers exports without complaint and comes out with rope on every layer.
Reproducing
examples/models/smollm3/3b_config.json and examples/models/smollm3/convert_weights.py are both in the tree, so SmolLM3-3B looks ready to export. Converting its weights and exporting through export_llm with that params json (XNNPACK, 8da4w + 8-bit embedding, 1852.2 MB) gives a .pte that loads and runs, then does this:
prompt: <|im_start|>user\nWhat is the capital of France?<|im_end|>\n<|im_start|>assistant\n
output: Okay, so you want to know what what what what what what what what what what what
what what what what what what what what what what what what what what what what
prompt: <|im_start|>user\nWhat is 17 times 4?<|im_end|>\n<|im_start|>assistant\n
output: I think you are looking for an answer that can be given that can be be be be be
determ determ determ determ determ determ determ determ determ determ#ae determ#
It starts as English and collapses, which is what a positional-encoding mismatch looks like: the first few tokens are fine and the drift grows with position.
What would help
Either implement the interval on this path, or reject a params json that sets a field the chosen backend does not implement. The silent version is the expensive one — the export succeeds, the file is the right size, and the model is simply a different model.
Measured with the executorch 1.4.0 wheel; the code paths quoted above are byte-identical on main at e4576d0.
cc @larryliu0820 @mergennachin @cccclai @helunwencser @jackzhxng @digantdesai
- Lenguaje dominante
- Python
- Estrellas
- 5.1k
- Forks
- 1.2k
- Merge medio
- 2 d 9 h
- PR fusionados (30 d)
- 661
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de pytorch/executorch
-
enhancement triaged
Dificultad 2/5 Medio día Aptitud para principiantes 68/100
pytorch/executorch#21640 ·
Los mantenedores suelen responder en 1 día
-
Prompts of max_seq_len tokens fail to prefill: export_llm bounds the KV-cache token input at max_seq_len - 1 but publishes get_max_seq_len = max_seq_lenPosiblemente ocupada @alpharomercoma la tomó hoy. Abierto
Dificultad 3/5 1-2 días Aptitud para principiantes 20/100
pytorch/executorch#23678 ·
Los mantenedores suelen responder en 1 día
-
[v1.6.0] Release Schedule and TrackerPosiblemente ocupada @JacobSzwejbka la tomó hace 2 días. Abierto
pytorch/executorch#23608 · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
QNN sharded LLM export fails: ResolveDebugHandle is not the last edge pass (dep_table entry overwritten by SplitGraph registration)Posiblemente ocupada @psiddh la tomó hace 2 días. Abiertobug module: llm module: qnn triaged
pytorch/executorch#23580 · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
Dificultad 4/5 3-5 días Aptitud para principiantes 48/100
pytorch/executorch#23553 ·
Los mantenedores suelen responder en 1 día
Todos los issues de pytorch/executorch
Issues similares
-
enhancement good first issue
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
-
Update Python support to 3.15Abiertopython-version
Dificultad 1/5 Menos de una hora Aptitud para principiantes 88/100
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
Los mantenedores suelen responder en 1 día
-
bug javascript P2-medium python release:v3.1
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
adrirubio/claude-deck#546 ·
Los mantenedores suelen responder en 1 día
-
area: desktop area: website priority: P2 type: feature
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
appandflow/stim#3411 · 1 comentario ·
Los mantenedores suelen responder en 1 día