MPS: scaled_softmax returns NaN — non_blocking=True host→MPS copy of the temperature tensor (decider-ai ≥ 1.4.0)
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 1/5
- Tempo stimato
- Meno di un'ora
- Idoneità per principianti
- 85/100
- Tipo di issue
- Bug
- Chiarezza
- Specificata chiaramente
- Stato di attività
- Attiva
- Ambito
- machine-learning
Direzione di ricerca
Il bug si trova in decider/temperature.py nel ramo delle liste della funzione scaled_softmax. Modifica la riga che crea il tensore di temperatura per rimuovere non_blocking=True o per costruirlo direttamente su lg.device. Conferma che la correzione funziona eseguendo lo script di riproduzione minimale fornito, che non dovrebbe più restituire NaN su MPS.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Summary
On Apple Silicon (MPS), every Decider.system_one call with decider-2b v11 (rev 533964da) returns NaN probabilities. In earlier runs it also gave a flat 1/3 or 1.000, varying from run to run. The model forward is correct; the fault is in the temperature step.
Cause
decider/temperature.py, scaled_softmax, list branch:
t = torch.tensor(temperature, dtype=lg.dtype).to(lg.device, non_blocking=True)[:, None]
v11's decider_config.json sets temperature_by_type, so every call takes this branch. The tensor is a temporary in pageable (unpinned) host memory; a non_blocking host→MPS copy from it can read the buffer before it is written or after it is freed, so the temperature arrives as garbage. On CUDA a copy from pageable memory is effectively synchronous, which is likely why it does not show up there. The by-type map arrived in decider-ai 1.4.0, so this looks like a regression from 1.4.0; the scalar-temperature path is unaffected.
Evidence
M5 Pro, macOS, torch 2.12.1, transformers 5.12.1, float32, the warm-up request (2+2? / 4 / 5, one choice question with A / B / TIE):
DecisionModelforward on MPS vs CPU: every decoder layer matches to < 4e-6 and the letter logits are identical (A 14.622, B 10.072, TIE 8.722), with and without the pad-to-64 and with and withoutattention_mask.scaled_softmax(lg, [1.164])on MPS, 500 calls: 500/500 NaN as shipped; 0/500 wrong with a synchronous.to(lg.device).system_one, 5 calls: NaN every time as shipped; A 0.9743 / B 0.0195 / TIE 0.0061 every time (identical to CPU) with only that copy made synchronous.
Minimal repro
import torch
from decider import temperature as TT
lg = torch.full((1, 16), float("-inf")); lg[0, :3] = torch.tensor([14.622, 10.072, 8.722])
print(TT.scaled_softmax(lg.to("mps"), [1.164])) # NaN on MPS
Suggested fix
Drop non_blocking=True, or build the tensor on the device directly:
t = torch.tensor(temperature, dtype=lg.dtype, device=lg.device)[:, None]
Workaround until then: Decider(..., temperature=<scalar>) turns the by-type map off, at the cost of the per-type calibration.
Separate note
decider.mps_ops.patch_mps() returns False on transformers < 5.17, so the MPS gated-delta patch silently does not run there. Not the cause of this bug, but a log line when the patch is skipped would help anyone diagnosing MPS behaviour.
- Lingua principale
- Python
- Stelle
- 1.1k
- Fork
- 46
- Merge medio
- 11h 12m
- PR unite (30g)
- 3
Preparare l'ambiente
Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di Mapika/decider
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
Tutte le issue di Mapika/decider
Issue simili
-
Difficoltà 1/5 1-3 ore Idoneità per principianti 85/100
pytest-dev/pluggy#757 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 1/5 1-3 ore Idoneità per principianti 85/100
NousResearch/hermes-agent#134960 ·
I maintainer di solito rispondono entro 1 giorno
-
HTML backend: `<br>` leaks the internal sentinel U+E000 into list items, headings and captionsForse già presa @morten-lagabote l’ha presa oggi. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 67/100
docling-project/docling#4671 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
I maintainer di solito rispondono entro 1 giorno
-
good first issue hacktoberfest infra
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
I maintainer di solito rispondono entro 1 giorno