Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Linux + RTX 50 (sm_120): an engine compiled with CUDA 13.2 answers garbage (IQ1_S/IQ2_S/IQ3_S miscompile); CUDA 13.0 works

Aperta
#892 0 commenti 1 reazione 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
48/100
Tipo di issue
Bug
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
cpp
Ambito
backend

Direzione di ricerca

Start with src/kernels/cuda/iq_kernels.cu, especially the IQ1_S, IQ2_S, and IQ3_S byte reads named in the report, and compare them with the referenced llama.cpp fix. Reproduce the CUDA 13.2 failure using the reported setup and, if available, run llama.cpp's test-backend-ops -b CUDA0. Done means the affected kernels pass and GPU inference returns correct answers rather than garbage.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

What happens

On Linux, setup finds no ready-made engine and compiles one with whatever toolkit is installed. With CUDA 13.2 on an RTX 5070 Ti, the engine loads and serves normally, but every answer is wrong:

  • "What is 2+2?" -> 4+4=8.8.
  • "Translate 'good morning' into German." -> Good morning
  • Most other prompts end after 1-2 tokens or loop (primary colors are primary colors are ...).

There are no errors in the log. The benchmark script's warm-up fails with no text received.

Environment

  • RTX 5070 Ti 16 GB (sm_120), driver 580.178.04, Ubuntu 24.04 (kernel 6.8.0-142), Core Ultra 7 265K (AVX2, no AVX-512), 62 GB RAM (KVM VM, GPU passthrough)
  • Strata 6f32ec0 (engine 0.1.39), ./setup.sh --yes --family qwen --model IQ2_XS --no-start
  • Setup compiled with /usr/local/cuda-13.2/bin/nvcc (13.2.51)
  • Model GGUF SHA-256s match the published LFS hashes (same as the RTX 5090 community report)

Isolation

Run Result
Strata engine, CUDA 13.2 (any --pcie-frac, --expert-cache 0, 4K or 64K context) garbage
llama.cpp 3cf0325 (Strata's pinned commit), CUDA 13.2, GPU offload, same GGUF garbage (0/12 on a small task suite)
llama.cpp, same build, -ngl 0 --device none (CPU only) correct
llama.cpp test-backend-ops -b CUDA0, CUDA 13.2 MUL_MAT 44 FAIL, MUL_MAT_ID 22 FAIL: only iq1_s, iq2_s, iq3_s (ERR 0.4-0.66); every other type OK
Strata engine rebuilt with CUDA 13.0.88 (pip nvidia-cuda-nvcc==13.0.88), nothing else changed correct: 12/12 on the task suite

This matches the known nvcc 13.2 / sm_120 miscompile of the byte reads in the IQ1_S/IQ2_S/IQ3_S kernels (ggml-org/llama.cpp#21255, #28581; the fix PR ggml-org/llama.cpp#28784 was closed unmerged). Strata's own src/kernels/cuda/iq_kernels.cu uses the same pattern (const uint8_t* qs = (const uint8_t*) &qs_packed; then qs[l] into iq2s_grid / iq3s_grid / iq1s_grid_gpu, e.g. lines 156, 215, 242, 658, 719, 753). I have not tested which of Strata's kernels and ggml's kernels are affected separately.

A PTX-only build (120-virtual) with 13.2 is not a workaround on driver 580: the provided PTX was compiled with an unsupported toolchain.

Results with the CUDA 13.0 engine (for reference)

bench/results/2026-09-30-community-rtx-5090/benchmark.py, 3 runs, 256 output tokens, 0 reused:

Prompt tokens Prompt tok/s Decode tok/s
4,096 2,706 128
32,768 3,280 130
60,000 3,196 125

Suggestions

  1. Setup: when the toolkit is nvcc 13.2 and a card is sm_120, warn or stop. Alternatively, compile with pip's nvidia-cuda-nvcc==13.0.88 (plus nvidia-cuda-cccl, nvidia-nvvm, nvidia-cuda-crt, and the runtime/cuBLAS wheels setup already pins). That works without sudo.
  2. Kernels: replace the uint8_t* byte indexing with __byte_perm / shifts in the IQ*_S paths (what ggml-org/llama.cpp#28784 did).
  3. A one-line self-check after the first start (for example, "capital of France" -> contains "Paris") would catch a broken build before users rely on it.
Lingua principale
C++
Stelle
11.6k
Fork
1k
Merge medio
7h 46m
PR unite (30g)
30

Preparare l'ambiente

Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di Niko1221/Strata

Tutte le issue di Niko1221/Strata

Issue simili

Altre issue su C++

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.