Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Linux + RTX 50 (sm_120): an engine compiled with CUDA 13.2 answers garbage (IQ1_S/IQ2_S/IQ3_S miscompile); CUDA 13.0 works

Open
#892 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
48/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
cpp
Domain
backend

Research direction

Start with src/kernels/cuda/iq_kernels.cu, especially the IQ1_S, IQ2_S, and IQ3_S byte reads named in the report, and compare them with the referenced llama.cpp fix. Reproduce the CUDA 13.2 failure using the reported setup and, if available, run llama.cpp's test-backend-ops -b CUDA0. Done means the affected kernels pass and GPU inference returns correct answers rather than garbage.

Written by the indexing model from the issue text.

Description

What happens

On Linux, setup finds no ready-made engine and compiles one with whatever toolkit is installed. With CUDA 13.2 on an RTX 5070 Ti, the engine loads and serves normally, but every answer is wrong:

  • "What is 2+2?" -> 4+4=8.8.
  • "Translate 'good morning' into German." -> Good morning
  • Most other prompts end after 1-2 tokens or loop (primary colors are primary colors are ...).

There are no errors in the log. The benchmark script's warm-up fails with no text received.

Environment

  • RTX 5070 Ti 16 GB (sm_120), driver 580.178.04, Ubuntu 24.04 (kernel 6.8.0-142), Core Ultra 7 265K (AVX2, no AVX-512), 62 GB RAM (KVM VM, GPU passthrough)
  • Strata 6f32ec0 (engine 0.1.39), ./setup.sh --yes --family qwen --model IQ2_XS --no-start
  • Setup compiled with /usr/local/cuda-13.2/bin/nvcc (13.2.51)
  • Model GGUF SHA-256s match the published LFS hashes (same as the RTX 5090 community report)

Isolation

Run Result
Strata engine, CUDA 13.2 (any --pcie-frac, --expert-cache 0, 4K or 64K context) garbage
llama.cpp 3cf0325 (Strata's pinned commit), CUDA 13.2, GPU offload, same GGUF garbage (0/12 on a small task suite)
llama.cpp, same build, -ngl 0 --device none (CPU only) correct
llama.cpp test-backend-ops -b CUDA0, CUDA 13.2 MUL_MAT 44 FAIL, MUL_MAT_ID 22 FAIL: only iq1_s, iq2_s, iq3_s (ERR 0.4-0.66); every other type OK
Strata engine rebuilt with CUDA 13.0.88 (pip nvidia-cuda-nvcc==13.0.88), nothing else changed correct: 12/12 on the task suite

This matches the known nvcc 13.2 / sm_120 miscompile of the byte reads in the IQ1_S/IQ2_S/IQ3_S kernels (ggml-org/llama.cpp#21255, #28581; the fix PR ggml-org/llama.cpp#28784 was closed unmerged). Strata's own src/kernels/cuda/iq_kernels.cu uses the same pattern (const uint8_t* qs = (const uint8_t*) &qs_packed; then qs[l] into iq2s_grid / iq3s_grid / iq1s_grid_gpu, e.g. lines 156, 215, 242, 658, 719, 753). I have not tested which of Strata's kernels and ggml's kernels are affected separately.

A PTX-only build (120-virtual) with 13.2 is not a workaround on driver 580: the provided PTX was compiled with an unsupported toolchain.

Results with the CUDA 13.0 engine (for reference)

bench/results/2026-09-30-community-rtx-5090/benchmark.py, 3 runs, 256 output tokens, 0 reused:

Prompt tokens Prompt tok/s Decode tok/s
4,096 2,706 128
32,768 3,280 130
60,000 3,196 125

Suggestions

  1. Setup: when the toolkit is nvcc 13.2 and a card is sm_120, warn or stop. Alternatively, compile with pip's nvidia-cuda-nvcc==13.0.88 (plus nvidia-cuda-cccl, nvidia-nvvm, nvidia-cuda-crt, and the runtime/cuBLAS wheels setup already pins). That works without sudo.
  2. Kernels: replace the uint8_t* byte indexing with __byte_perm / shifts in the IQ*_S paths (what ggml-org/llama.cpp#28784 did).
  3. A one-line self-check after the first start (for example, "capital of France" -> contains "Paris") would catch a broken build before users rely on it.
Dominant language
C++
Stars
11.6k
Forks
1k
Avg merge
7h 46m
Merged PRs (30d)
30

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Niko1221/Strata

All issues in Niko1221/Strata

Similar issues

More C++ issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.