Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

MPS: scaled_softmax returns NaN — non_blocking=True host→MPS copy of the temperature tensor (decider-ai ≥ 1.4.0)

Open Beginner friendly
#21 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
1/5
Estimated time
Under an hour
Newbie friendliness
85/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python, pytorch

Research direction

The bug is in decider/temperature.py in the scaled_softmax function's list branch. Modify the line that creates the temperature tensor to either remove non_blocking=True or build it directly on lg.device. Confirm the fix works by running the provided minimal reproduction script, which should no longer return NaN on MPS.

Written by the indexing model from the issue text.

Description

Summary

On Apple Silicon (MPS), every Decider.system_one call with decider-2b v11 (rev 533964da) returns NaN probabilities. In earlier runs it also gave a flat 1/3 or 1.000, varying from run to run. The model forward is correct; the fault is in the temperature step.

Cause

decider/temperature.py, scaled_softmax, list branch:

t = torch.tensor(temperature, dtype=lg.dtype).to(lg.device, non_blocking=True)[:, None]

v11's decider_config.json sets temperature_by_type, so every call takes this branch. The tensor is a temporary in pageable (unpinned) host memory; a non_blocking host→MPS copy from it can read the buffer before it is written or after it is freed, so the temperature arrives as garbage. On CUDA a copy from pageable memory is effectively synchronous, which is likely why it does not show up there. The by-type map arrived in decider-ai 1.4.0, so this looks like a regression from 1.4.0; the scalar-temperature path is unaffected.

Evidence

M5 Pro, macOS, torch 2.12.1, transformers 5.12.1, float32, the warm-up request (2+2? / 4 / 5, one choice question with A / B / TIE):

  • DecisionModel forward on MPS vs CPU: every decoder layer matches to < 4e-6 and the letter logits are identical (A 14.622, B 10.072, TIE 8.722), with and without the pad-to-64 and with and without attention_mask.
  • scaled_softmax(lg, [1.164]) on MPS, 500 calls: 500/500 NaN as shipped; 0/500 wrong with a synchronous .to(lg.device).
  • system_one, 5 calls: NaN every time as shipped; A 0.9743 / B 0.0195 / TIE 0.0061 every time (identical to CPU) with only that copy made synchronous.

Minimal repro

import torch
from decider import temperature as TT
lg = torch.full((1, 16), float("-inf")); lg[0, :3] = torch.tensor([14.622, 10.072, 8.722])
print(TT.scaled_softmax(lg.to("mps"), [1.164]))   # NaN on MPS

Suggested fix

Drop non_blocking=True, or build the tensor on the device directly:

t = torch.tensor(temperature, dtype=lg.dtype, device=lg.device)[:, None]

Workaround until then: Decider(..., temperature=<scalar>) turns the by-type map off, at the cost of the per-type calibration.

Separate note

decider.mps_ops.patch_mps() returns False on transformers < 5.17, so the MPS gated-delta patch silently does not run there. Not the cause of this bug, but a log line when the patch is skipped would help anyone diagnosing MPS behaviour.

Dominant language
Python
Stars
1.1k
Forks
46
Avg merge
11h 12m
Merged PRs (30d)
3

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Mapika/decider

All issues in Mapika/decider

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.