Linux + RTX 50 (sm_120): an engine compiled with CUDA 13.2 answers garbage (IQ1_S/IQ2_S/IQ3_S miscompile); CUDA 13.0 works
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 48/100
Research direction
Start with src/kernels/cuda/iq_kernels.cu, especially the IQ1_S, IQ2_S, and IQ3_S byte reads named in the report, and compare them with the referenced llama.cpp fix. Reproduce the CUDA 13.2 failure using the reported setup and, if available, run llama.cpp's test-backend-ops -b CUDA0. Done means the affected kernels pass and GPU inference returns correct answers rather than garbage.
Written by the indexing model from the issue text.
Description
What happens
On Linux, setup finds no ready-made engine and compiles one with whatever toolkit is installed. With CUDA 13.2 on an RTX 5070 Ti, the engine loads and serves normally, but every answer is wrong:
- "What is 2+2?" ->
4+4=8.8. - "Translate 'good morning' into German." ->
Good morning - Most other prompts end after 1-2 tokens or loop (
primary colors are primary colors are ...).
There are no errors in the log. The benchmark script's warm-up fails with no text received.
Environment
- RTX 5070 Ti 16 GB (sm_120), driver 580.178.04, Ubuntu 24.04 (kernel 6.8.0-142), Core Ultra 7 265K (AVX2, no AVX-512), 62 GB RAM (KVM VM, GPU passthrough)
- Strata
6f32ec0(engine 0.1.39),./setup.sh --yes --family qwen --model IQ2_XS --no-start - Setup compiled with
/usr/local/cuda-13.2/bin/nvcc(13.2.51) - Model GGUF SHA-256s match the published LFS hashes (same as the RTX 5090 community report)
Isolation
| Run | Result |
|---|---|
Strata engine, CUDA 13.2 (any --pcie-frac, --expert-cache 0, 4K or 64K context) |
garbage |
llama.cpp 3cf0325 (Strata's pinned commit), CUDA 13.2, GPU offload, same GGUF |
garbage (0/12 on a small task suite) |
llama.cpp, same build, -ngl 0 --device none (CPU only) |
correct |
llama.cpp test-backend-ops -b CUDA0, CUDA 13.2 |
MUL_MAT 44 FAIL, MUL_MAT_ID 22 FAIL: only iq1_s, iq2_s, iq3_s (ERR 0.4-0.66); every other type OK |
Strata engine rebuilt with CUDA 13.0.88 (pip nvidia-cuda-nvcc==13.0.88), nothing else changed |
correct: 12/12 on the task suite |
This matches the known nvcc 13.2 / sm_120 miscompile of the byte reads in the IQ1_S/IQ2_S/IQ3_S kernels (ggml-org/llama.cpp#21255, #28581; the fix PR ggml-org/llama.cpp#28784 was closed unmerged). Strata's own src/kernels/cuda/iq_kernels.cu uses the same pattern (const uint8_t* qs = (const uint8_t*) &qs_packed; then qs[l] into iq2s_grid / iq3s_grid / iq1s_grid_gpu, e.g. lines 156, 215, 242, 658, 719, 753). I have not tested which of Strata's kernels and ggml's kernels are affected separately.
A PTX-only build (120-virtual) with 13.2 is not a workaround on driver 580: the provided PTX was compiled with an unsupported toolchain.
Results with the CUDA 13.0 engine (for reference)
bench/results/2026-09-30-community-rtx-5090/benchmark.py, 3 runs, 256 output tokens, 0 reused:
| Prompt tokens | Prompt tok/s | Decode tok/s |
|---|---|---|
| 4,096 | 2,706 | 128 |
| 32,768 | 3,280 | 130 |
| 60,000 | 3,196 | 125 |
Suggestions
- Setup: when the toolkit is nvcc 13.2 and a card is sm_120, warn or stop. Alternatively, compile with pip's
nvidia-cuda-nvcc==13.0.88(plusnvidia-cuda-cccl,nvidia-nvvm,nvidia-cuda-crt, and the runtime/cuBLAS wheels setup already pins). That works without sudo. - Kernels: replace the
uint8_t*byte indexing with__byte_perm/ shifts in the IQ*_S paths (what ggml-org/llama.cpp#28784 did). - A one-line self-check after the first start (for example, "capital of France" -> contains "Paris") would catch a broken build before users rely on it.
- Dominant language
- C++
- Stars
- 11.6k
- Forks
- 1k
- Avg merge
- 7h 46m
- Merged PRs (30d)
- 30
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Niko1221/Strata
-
ROCm on Windows ZIPOpen
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
Niko1221/Strata#738 · 1 comment · 1 reaction ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 67/100
Niko1221/Strata#694 · 2 comments · 4 reactions ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Niko1221/Strata#691 · 2 comments ·
Maintainers usually reply within 1 day
Similar issues
-
enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
iOS: hidden scale bar invalidates its intrinsic content size on every layout pass of MLNMapViewOpen
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
maplibre/maplibre-native#4723 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
HarbourMasters/Shipwright#7320 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
-
Make Catch2 optional when `RDK_BUILD_CPP_TESTS=OFF`Possibly taken @pechersky claimed this today. Openbug
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
Maintainers usually reply within 2 days