Linux + RTX 50 (sm_120): an engine compiled with CUDA 13.2 answers garbage (IQ1_S/IQ2_S/IQ3_S miscompile); CUDA 13.0 works
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 48/100
Direzione di ricerca
Start with src/kernels/cuda/iq_kernels.cu, especially the IQ1_S, IQ2_S, and IQ3_S byte reads named in the report, and compare them with the referenced llama.cpp fix. Reproduce the CUDA 13.2 failure using the reported setup and, if available, run llama.cpp's test-backend-ops -b CUDA0. Done means the affected kernels pass and GPU inference returns correct answers rather than garbage.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
What happens
On Linux, setup finds no ready-made engine and compiles one with whatever toolkit is installed. With CUDA 13.2 on an RTX 5070 Ti, the engine loads and serves normally, but every answer is wrong:
- "What is 2+2?" ->
4+4=8.8. - "Translate 'good morning' into German." ->
Good morning - Most other prompts end after 1-2 tokens or loop (
primary colors are primary colors are ...).
There are no errors in the log. The benchmark script's warm-up fails with no text received.
Environment
- RTX 5070 Ti 16 GB (sm_120), driver 580.178.04, Ubuntu 24.04 (kernel 6.8.0-142), Core Ultra 7 265K (AVX2, no AVX-512), 62 GB RAM (KVM VM, GPU passthrough)
- Strata
6f32ec0(engine 0.1.39),./setup.sh --yes --family qwen --model IQ2_XS --no-start - Setup compiled with
/usr/local/cuda-13.2/bin/nvcc(13.2.51) - Model GGUF SHA-256s match the published LFS hashes (same as the RTX 5090 community report)
Isolation
| Run | Result |
|---|---|
Strata engine, CUDA 13.2 (any --pcie-frac, --expert-cache 0, 4K or 64K context) |
garbage |
llama.cpp 3cf0325 (Strata's pinned commit), CUDA 13.2, GPU offload, same GGUF |
garbage (0/12 on a small task suite) |
llama.cpp, same build, -ngl 0 --device none (CPU only) |
correct |
llama.cpp test-backend-ops -b CUDA0, CUDA 13.2 |
MUL_MAT 44 FAIL, MUL_MAT_ID 22 FAIL: only iq1_s, iq2_s, iq3_s (ERR 0.4-0.66); every other type OK |
Strata engine rebuilt with CUDA 13.0.88 (pip nvidia-cuda-nvcc==13.0.88), nothing else changed |
correct: 12/12 on the task suite |
This matches the known nvcc 13.2 / sm_120 miscompile of the byte reads in the IQ1_S/IQ2_S/IQ3_S kernels (ggml-org/llama.cpp#21255, #28581; the fix PR ggml-org/llama.cpp#28784 was closed unmerged). Strata's own src/kernels/cuda/iq_kernels.cu uses the same pattern (const uint8_t* qs = (const uint8_t*) &qs_packed; then qs[l] into iq2s_grid / iq3s_grid / iq1s_grid_gpu, e.g. lines 156, 215, 242, 658, 719, 753). I have not tested which of Strata's kernels and ggml's kernels are affected separately.
A PTX-only build (120-virtual) with 13.2 is not a workaround on driver 580: the provided PTX was compiled with an unsupported toolchain.
Results with the CUDA 13.0 engine (for reference)
bench/results/2026-09-30-community-rtx-5090/benchmark.py, 3 runs, 256 output tokens, 0 reused:
| Prompt tokens | Prompt tok/s | Decode tok/s |
|---|---|---|
| 4,096 | 2,706 | 128 |
| 32,768 | 3,280 | 130 |
| 60,000 | 3,196 | 125 |
Suggestions
- Setup: when the toolkit is nvcc 13.2 and a card is sm_120, warn or stop. Alternatively, compile with pip's
nvidia-cuda-nvcc==13.0.88(plusnvidia-cuda-cccl,nvidia-nvvm,nvidia-cuda-crt, and the runtime/cuBLAS wheels setup already pins). That works without sudo. - Kernels: replace the
uint8_t*byte indexing with__byte_perm/ shifts in the IQ*_S paths (what ggml-org/llama.cpp#28784 did). - A one-line self-check after the first start (for example, "capital of France" -> contains "Paris") would catch a broken build before users rely on it.
- Lingua principale
- C++
- Stelle
- 11.6k
- Fork
- 1k
- Merge medio
- 7h 46m
- PR unite (30g)
- 30
Preparare l'ambiente
Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di Niko1221/Strata
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
Niko1221/Strata#738 · 1 commento · 1 reazione ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 67/100
Niko1221/Strata#694 · 2 commenti · 4 reazioni ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
Niko1221/Strata#691 · 2 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 80/100
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di Niko1221/Strata
Issue simili
-
ws_bridge: stripping format=evr for matchmaker connections can concatenate the path and queryApertabug
Difficoltà 2/5 1-3 ore Idoneità per principianti 73/100
EchoTools/nevr-runtime#116 ·
I maintainer di solito rispondono entro 1 giorno
-
code-quality libc++
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 82/100
llvm/llvm-project#229284 ·
I maintainer di solito rispondono entro 1 giorno
-
test-issue
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
llvm/offload-test-suite#1557 ·
I maintainer di solito rispondono entro 1 giorno
-
enhancement
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno
-
iOS: hidden scale bar invalidates its intrinsic content size on every layout pass of MLNMapViewAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
maplibre/maplibre-native#4723 ·
I maintainer di solito rispondono entro 1 giorno