Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Windows on ARM (Snapdragon X): numpy<2 pin blocks install; bf16 on CPU is 13× slower; llama.cpp needs clang; QNN graph-size limit

Open
#18 6 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
35/100
Issue type
Feature
Clarity
Needs clarification
Activity status
Active
Tech stack
python, pytorch, shell

Research direction

Start by separating the five findings and reading pyproject.toml, engine.py, scripts/serve.sh, and the referenced QNN diagnosis scripts decider/qnn_bisect.py and decider/qnn_chain.py. Use the Snapdragon findings and docs/FINDINGS.md §3–4 for platform context. Done should mean each selected item has either an implemented change with coverage or clear documentation, with install, CPU, serving, and QNN behavior verified.

Written by the indexing model from the issue text.

Description

Notes from getting decider running on a Snapdragon X Elite laptop (Windows 11 ARM64, native CPython 3.12, PyTorch 2.14.0+cpu, decider-ai 1.5.0, decider-2b v11). It works, and the GGUF path is fast. Five things tripped us up; any of them could be fixed or just documented.

1. numpy<2 pin blocks install on Windows ARM64

pyproject.toml pins numpy<2. numpy 1.26 has no win_arm64 cp312 wheel, so pip tries (and fails) to build it from source. The only runtime numpy use we found (np.full in engine.py) works on numpy 2. We installed with --no-deps on top of numpy 2 and everything ran. Could the pin be relaxed?

2. bfloat16 on CPU is about 13× slower than float32 here

The CPU path defaults to bfloat16. On Windows ARM64 PyTorch 2.14 there's no fast bf16 kernel path: decider-2b took 137 s for a 3-question request in bf16 vs 11 s in float32. Suggest defaulting to float32 on CPU (or when bf16 isn't accelerated), or documenting dtype=torch.float32.

3. GGUF via llama-cpp-python needs clang, and CPU-only was the stable build

  • llama-cpp-python 0.3.35 won't build with MSVC on ARM64 (ggml-cpu: "MSVC is not supported for ARM, use clang"). With the Visual Studio Build Tools "C++ Clang tools for Windows" component it builds with -DGGML_OPENMP=OFF and every GPU backend off.
  • Result: decider-2b v11 Q8_0 at about 220 ms per 3-question request on 8 threads (Q4_K_M about 330 ms; Q8_0 is faster on this CPU). That is about 50× faster than PyTorch float32.
  • Caution: the laptop blue-screened twice (0xD1) while we tried prebuilt llama.cpp with the Adreno OpenCL backend. The faulting driver isn't confirmed yet, but CPU-only was stable throughout.

4. Qualcomm GPU (QNN) export hits a per-process graph-size limit

A static ONNX export of decider-2b (S=256, fp16) unrolls the Gated DeltaNet layers into about 20,000 nodes (about 2,350 Scatter ops). ONNX Runtime CPU matches the GGUF answers. onnxruntime-qnn 2.6.0's GPU backend compiles any piece up to about 3,000 nodes, but caps the cumulative total per process, so the model can't run there as one graph or as chained segments. A compact DeltaNet export (a recurrent form instead of the unrolled chunk scan) would probably be needed for Qualcomm accelerators. Diagnosis scripts: decider/qnn_bisect.py and decider/qnn_chain.py in the repo below.

5. scripts/serve.sh binds 0.0.0.0 with no auth

--host 0.0.0.0 exposes the server to the local network by default. Consider 127.0.0.1 as the default, with an opt-in for 0.0.0.0.

Setup scripts, the clang build recipe and full numbers: https://github.com/esterhuizen/system-one-on-snapdragon (docs/FINDINGS.md §3–4). Happy to turn 1, 2 and 5 into a PR if you'd like.

Dominant language
Python
Stars
1.1k
Forks
46
Avg merge
14h 48m
Merged PRs (30d)
1

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Mapika/decider

All issues in Mapika/decider

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.