Windows on ARM (Snapdragon X): numpy<2 pin blocks install; bf16 on CPU is 13× slower; llama.cpp needs clang; QNN graph-size limit
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 35/100
- Issue type
- Feature
- Clarity
- Needs clarification
- Activity status
- Active
Research direction
Start by separating the five findings and reading pyproject.toml, engine.py, scripts/serve.sh, and the referenced QNN diagnosis scripts decider/qnn_bisect.py and decider/qnn_chain.py. Use the Snapdragon findings and docs/FINDINGS.md §3–4 for platform context. Done should mean each selected item has either an implemented change with coverage or clear documentation, with install, CPU, serving, and QNN behavior verified.
Written by the indexing model from the issue text.
Description
Notes from getting decider running on a Snapdragon X Elite laptop (Windows 11 ARM64, native CPython 3.12, PyTorch 2.14.0+cpu, decider-ai 1.5.0, decider-2b v11). It works, and the GGUF path is fast. Five things tripped us up; any of them could be fixed or just documented.
1. numpy<2 pin blocks install on Windows ARM64
pyproject.toml pins numpy<2. numpy 1.26 has no win_arm64 cp312 wheel, so pip tries (and fails) to build it from source. The only runtime numpy use we found (np.full in engine.py) works on numpy 2. We installed with --no-deps on top of numpy 2 and everything ran. Could the pin be relaxed?
2. bfloat16 on CPU is about 13× slower than float32 here
The CPU path defaults to bfloat16. On Windows ARM64 PyTorch 2.14 there's no fast bf16 kernel path: decider-2b took 137 s for a 3-question request in bf16 vs 11 s in float32. Suggest defaulting to float32 on CPU (or when bf16 isn't accelerated), or documenting dtype=torch.float32.
3. GGUF via llama-cpp-python needs clang, and CPU-only was the stable build
- llama-cpp-python 0.3.35 won't build with MSVC on ARM64 (ggml-cpu: "MSVC is not supported for ARM, use clang"). With the Visual Studio Build Tools "C++ Clang tools for Windows" component it builds with
-DGGML_OPENMP=OFFand every GPU backend off. - Result: decider-2b v11 Q8_0 at about 220 ms per 3-question request on 8 threads (Q4_K_M about 330 ms; Q8_0 is faster on this CPU). That is about 50× faster than PyTorch float32.
- Caution: the laptop blue-screened twice (0xD1) while we tried prebuilt llama.cpp with the Adreno OpenCL backend. The faulting driver isn't confirmed yet, but CPU-only was stable throughout.
4. Qualcomm GPU (QNN) export hits a per-process graph-size limit
A static ONNX export of decider-2b (S=256, fp16) unrolls the Gated DeltaNet layers into about 20,000 nodes (about 2,350 Scatter ops). ONNX Runtime CPU matches the GGUF answers. onnxruntime-qnn 2.6.0's GPU backend compiles any piece up to about 3,000 nodes, but caps the cumulative total per process, so the model can't run there as one graph or as chained segments. A compact DeltaNet export (a recurrent form instead of the unrolled chunk scan) would probably be needed for Qualcomm accelerators. Diagnosis scripts: decider/qnn_bisect.py and decider/qnn_chain.py in the repo below.
5. scripts/serve.sh binds 0.0.0.0 with no auth
--host 0.0.0.0 exposes the server to the local network by default. Consider 127.0.0.1 as the default, with an opt-in for 0.0.0.0.
Setup scripts, the clang build recipe and full numbers: https://github.com/esterhuizen/system-one-on-snapdragon (docs/FINDINGS.md §3–4). Happy to turn 1, 2 and 5 into a PR if you'd like.
- Dominant language
- Python
- Stars
- 1.1k
- Forks
- 46
- Avg merge
- 14h 48m
- Merged PRs (30d)
- 1
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Mapika/decider
-
Difficulty 1/5 Under an hour Newbie friendliness 1/100
-
Difficulty 5/5 Over a week Newbie friendliness 35/100
Similar issues
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
Deepak3699/Ai_Mentor#244 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
btclib-org/btclib-node#1880 ·
Maintainers usually reply within 1 day
-
CONTRIBUTING.md: say how ticketless bug fixes and feature PRs are handledPossibly taken @khuisman claimed this today. Openv0.9.2
Difficulty 1/5 Under an hour Newbie friendliness 84/100
khuisman/mcp-gee-sweet#941 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
pyjanitor-devs/pyjanitor#1758 ·
Maintainers usually reply within 1 day
-
bug ready for review
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
odysseus-dev/odysseus#6641 ·
Maintainers usually reply within 1 day