Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Fresh prompts over --short-read decode !!!! (token 0) when the GPU holds 100% of the experts (48 GB card)

Open
#871 1 comment 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
35/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
cpp

Research direction

Start by comparing the all-resident (“zero-doorbell”) batched prompt-read path referenced in #646 with the related report in #792; BATCHING.md is also mentioned, though its measurements exclude 100% residency. Reproduce with a fresh prompt longer than --short-read and the expert cache at full residency, then compare against a cache below full residency. Done means fresh prompts produce normal output when all experts are resident, without breaking the working partial-residency path.

Written by the indexing model from the issue text.

Description

Fresh prompts over --short-read decode !!!! (token 0) when the GPU holds 100% of the experts (48 GB card)

Summary

On a card large enough to hold every expert of IQ2_XS in VRAM, any fresh prompt whose new text exceeds
--short-read (default 64 tokens) answers a run of ! (token 0). Shorter prompts, and requests that hit the
prompt cache, answer normally. Capping the expert cache below 100% residency makes it go away with no other
change. This looks like the all-resident ("zero-doorbell", #646) path on the batched prompt read, the same
family as #792, which notes that BATCHING.md was measured only on cards below 100% residency. A 24 GB card never
reaches full residency with IQ2_XS, which would explain why this has not been reported.

Environment

  • Strata 6f32ec0, ready-made engine 0.1.39 (sm_75/86/89/120, CUDA 13.0), Windows 11 Pro 26200
  • GPU: RTX 4090 48 GB (memory-modded board), compute capability 8.9, driver 591.86
  • CPU: Intel i7-14700K (AVX2, no AVX-512, 8P+12E, --pool-workers 13 chosen by setup); RAM: 64 GB
  • Model: Qwen3.8-Flash-Next GSQ-RCO IQ2_XS (original family)

Reproduction (untouched install)

  1. Fresh clone, START-HERE.bat --yes (no other flags). Setup's settings:
    --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --max-context 131072 --kv int8 --kv-resident 32768 --pool-workers 13
  2. Startup log: filling the GPU's expert cache (24576 experts, 33.02 GiB of VRAM) /
    expert cache 24576 slots, 33.02 GiB of VRAM
  3. Send a fresh ~350-token user message (any content; a random prefix avoids the prompt cache),
    reasoning_effort: "none".

Result: 4/4 requests answer !!!!!!!!…. Sampling makes no difference (temperature 0, 0.7, absent).
The serve log for the bad requests shows every draft accepted:

strata serve: prompt 354 tokens = 0 reused + 354 read in 390 ms (907.6 tok/s), 40 generated in 219 ms (182.8 tok/s), drafts accepted 18 of 18, 1 checkpoints

What narrows it down

Threshold is exactly --short-read. Same prompt truncated word by word, fresh each time:

prompt tokens (incl. template) result
28, 35, 70 ok
72, 75, 76, 80, 92, 137, 237, 354 !!!!

The decode path is fine.

  • --short-read 1000000 (every fresh part through the decode windows): 0/3 bad, but a 40K prompt reads at 352 tok/s.
  • Repeating the same prompt (prompt cache hit, only the tail read through the decode path) answers correctly, even
    though the reused KV came from the batched read. So the KV written by the batched read looks fine; what breaks
    is the handoff from the batched read to the first decoded token.

It depends on residency, not on anything else. Same config, only --expert-cache changed, no other workaround:

--expert-cache slots the engine allocated fresh 354-token prompt
auto 24576 (100%) 4/4 bad
24000 24576 (rounded up to 100%) 4/4 bad
23000 24076 (98%) 0/4 bad
20000 20939 (85%) 0/4 bad

With 23000: 40K prompt 3,780 tok/s, decode 195–217 tok/s, which is what we run now.

Ruled out (still bad with each): low-RAM resident mode, --low-ram mmap, normal RAM mode; --vision on/off;
KV streaming on/off (--max-context 32768, no --kv-resident); --spec 2; --pcie-frac 0;
STRATA_RESIDENT_PIN=0, STRATA_ARGMAX_MULTI=0, STRATA_PF_FUSED=0, STRATA_PLE_BATCH=0.
CUDA_LAUNCH_BLOCKING=1 makes the engine exit with 0xC0000409 (stall dump written).

Two smaller points

  • --expert-cache 24000 rounds up to the full 24,576, so a user picking a "just under" value to avoid the
    all-resident path lands back on it. Some note or a hard cap below the total would help.
  • With a 48 GB card the default auto always fills the cache, so every fresh prompt over 64 tokens fails out of the
    box. Most clients (agents with a system prompt plus tool schemas) never send one under 64 tokens.

Happy to run a debug build or STRATA_* switches on this machine if useful.

Dominant language
C++
Stars
11.6k
Forks
1k
Avg merge
7h 46m
Merged PRs (30d)
30

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Niko1221/Strata

All issues in Niko1221/Strata

Similar issues

More C++ issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.