Fresh prompts over --short-read decode !!!! (token 0) when the GPU holds 100% of the experts (48 GB card)
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 35/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- cpp
- Domain
- machine-learning
Research direction
Start by comparing the all-resident (“zero-doorbell”) batched prompt-read path referenced in #646 with the related report in #792; BATCHING.md is also mentioned, though its measurements exclude 100% residency. Reproduce with a fresh prompt longer than --short-read and the expert cache at full residency, then compare against a cache below full residency. Done means fresh prompts produce normal output when all experts are resident, without breaking the working partial-residency path.
Written by the indexing model from the issue text.
Description
Fresh prompts over --short-read decode !!!! (token 0) when the GPU holds 100% of the experts (48 GB card)
Summary
On a card large enough to hold every expert of IQ2_XS in VRAM, any fresh prompt whose new text exceeds
--short-read (default 64 tokens) answers a run of ! (token 0). Shorter prompts, and requests that hit the
prompt cache, answer normally. Capping the expert cache below 100% residency makes it go away with no other
change. This looks like the all-resident ("zero-doorbell", #646) path on the batched prompt read, the same
family as #792, which notes that BATCHING.md was measured only on cards below 100% residency. A 24 GB card never
reaches full residency with IQ2_XS, which would explain why this has not been reported.
Environment
- Strata
6f32ec0, ready-made engine 0.1.39 (sm_75/86/89/120, CUDA 13.0), Windows 11 Pro 26200 - GPU: RTX 4090 48 GB (memory-modded board), compute capability 8.9, driver 591.86
- CPU: Intel i7-14700K (AVX2, no AVX-512, 8P+12E,
--pool-workers 13chosen by setup); RAM: 64 GB - Model: Qwen3.8-Flash-Next GSQ-RCO IQ2_XS (original family)
Reproduction (untouched install)
- Fresh clone,
START-HERE.bat --yes(no other flags). Setup's settings:
--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --max-context 131072 --kv int8 --kv-resident 32768 --pool-workers 13 - Startup log:
filling the GPU's expert cache (24576 experts, 33.02 GiB of VRAM)/
expert cache 24576 slots, 33.02 GiB of VRAM - Send a fresh ~350-token user message (any content; a random prefix avoids the prompt cache),
reasoning_effort: "none".
Result: 4/4 requests answer !!!!!!!!…. Sampling makes no difference (temperature 0, 0.7, absent).
The serve log for the bad requests shows every draft accepted:
strata serve: prompt 354 tokens = 0 reused + 354 read in 390 ms (907.6 tok/s), 40 generated in 219 ms (182.8 tok/s), drafts accepted 18 of 18, 1 checkpoints
What narrows it down
Threshold is exactly --short-read. Same prompt truncated word by word, fresh each time:
| prompt tokens (incl. template) | result |
|---|---|
| 28, 35, 70 | ok |
| 72, 75, 76, 80, 92, 137, 237, 354 | !!!! |
The decode path is fine.
--short-read 1000000(every fresh part through the decode windows): 0/3 bad, but a 40K prompt reads at 352 tok/s.- Repeating the same prompt (prompt cache hit, only the tail read through the decode path) answers correctly, even
though the reused KV came from the batched read. So the KV written by the batched read looks fine; what breaks
is the handoff from the batched read to the first decoded token.
It depends on residency, not on anything else. Same config, only --expert-cache changed, no other workaround:
--expert-cache |
slots the engine allocated | fresh 354-token prompt |
|---|---|---|
| auto | 24576 (100%) | 4/4 bad |
| 24000 | 24576 (rounded up to 100%) | 4/4 bad |
| 23000 | 24076 (98%) | 0/4 bad |
| 20000 | 20939 (85%) | 0/4 bad |
With 23000: 40K prompt 3,780 tok/s, decode 195–217 tok/s, which is what we run now.
Ruled out (still bad with each): low-RAM resident mode, --low-ram mmap, normal RAM mode; --vision on/off;
KV streaming on/off (--max-context 32768, no --kv-resident); --spec 2; --pcie-frac 0;
STRATA_RESIDENT_PIN=0, STRATA_ARGMAX_MULTI=0, STRATA_PF_FUSED=0, STRATA_PLE_BATCH=0.
CUDA_LAUNCH_BLOCKING=1 makes the engine exit with 0xC0000409 (stall dump written).
Two smaller points
--expert-cache 24000rounds up to the full 24,576, so a user picking a "just under" value to avoid the
all-resident path lands back on it. Some note or a hard cap below the total would help.- With a 48 GB card the default
autoalways fills the cache, so every fresh prompt over 64 tokens fails out of the
box. Most clients (agents with a system prompt plus tool schemas) never send one under 64 tokens.
Happy to run a debug build or STRATA_* switches on this machine if useful.
- Dominant language
- C++
- Stars
- 11.6k
- Forks
- 1k
- Avg merge
- 7h 46m
- Merged PRs (30d)
- 30
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Niko1221/Strata
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Niko1221/Strata#974 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
Similar issues
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 73/100
EchoTools/nevr-runtime#116 ·
Maintainers usually reply within 1 day
-
code-quality libc++
Difficulty 1/5 Under an hour Newbie friendliness 82/100
llvm/llvm-project#229284 ·
Maintainers usually reply within 1 day
-
test-issue
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
llvm/offload-test-suite#1557 ·
Maintainers usually reply within 1 day
-
enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
iOS: hidden scale bar invalidates its intrinsic content size on every layout pass of MLNMapViewOpen
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
maplibre/maplibre-native#4723 ·
Maintainers usually reply within 1 day