HIP gfx1030, 2x RX 6900 XT: verify hangs after a long prompt when MMQ prompt path, adaptive expert swaps and SDMA meet (any one off avoids it)
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 25/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- cpp, linux
- Domain
- machine-learning
Research direction
Start by reproducing the hang with the request sequence and settings in the issue, then read the HSA queue head packets and signal values from inside the process through an HSA_TOOLS_LIB shim. Compare those signals with the reported cross-queue wait to determine whether it has completed or is still waiting on a dependency. Done means identifying the cause of the hang; the issue does not name specific source files or tests.
Written by the indexing model from the issue text.
Description
Summary
On 2x RX 6900 XT (gfx1030, HIP, Linux), engine 0.1.39 hangs on the first verify window after a ~31K-token prompt, once the engine has served about 7 earlier requests: verify: timed out at layer 1; its GPU waits were released but the GPU did not finish within 5 s (#267). Stock 6f32ec0 at default settings hung in 5 of 5 runs of the sequence below (and in 11 of 12 more runs of it with one diagnostic setting that did not avoid it, listed further down).
It needs three things at once; taking away any one of them stops it in my runs: the MMQ prompt path (STRATA_PREFILL_MMQ=0: 0/4), the primary card's adaptive expert swaps (--adapt-every 100000: 0/7), and SDMA copies (HSA_ENABLE_SDMA=0: 0/4). At the hang the primary card is 98-99% busy with no waves resident: its command processor is stuck at a wait, not a kernel. The root cause is not known yet; what is measured and what is not is listed below.
This is not #649, although the message is the same; see the last section.
Environment
- Strata
6f32ec0(engine 0.1.39),build-hipwith-DSTRATA_PREFILL_MMQ=ON(as setup.py builds for AMD), gfx1030 - 2x RX 6900 XT 16 GB, PCIe 4.0 x8 each; Ryzen 5 5600X; 128 GB DDR4 (about 69 GB still free at the hang)
- Linux 7.0, ROCm 10.0
- Model: Qwen3.8-Flash-Next GSQ-RCO IQ3_S, native pack,
--ple-gguf,--mtpdraft head - Arguments:
--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --max-context 131072 --kv int8 --kv-resident 32768 --expert-cache-device1 auto --remote-expert-opt --pcie-frac 0,HIP_VISIBLE_DEVICES=1,0(the expert arena, 46.84 GiB, is loaded whole andcudaHostRegistered)
Reproduction
A fresh server, then 8 chat requests, one after the other (greedy, thinking off, max_tokens 256): 7 short ones of 1.0K-4.9K tokens, then one of 31,219 tokens. The texts are slices of the repository's own docs and sources; the content does not matter (other slices hang the same way).
- Stock, default settings, this sequence: 5 of 5 runs hung on the 31K request. With
--prompt-cache 12instead of the default 6: 1 of 2. - With 5 or fewer earlier requests before the 31K one: 0 of 5 (one of those with
--prompt-cache 2). - Every stock hang was the first verify window after the batched prompt (6 tokens at position 31212). The one hang in the batched prompt read itself -
no progress for 60 s ... (reading the prompt (batched): waiting for the GPU (attention, router) at layer 1 of the prompt chunk from token 0)- was in theSTRATA_SH_STREAM=0arm below.
STRATA_VERIFY_TRACE=1
The host serves layer 0 within about 3.3 ms (flag 1 A 1 B 1). The GPU's breadcrumbs stop after layer 0's waitCPU; combine is never stamped:
layer 0 group 0: pre(0)=0.0 hc-read0(1)=101.6 ... waitB(21)=661.2 PCIe grp(22)=672.4 waitCPU(23)=3849.1
TIMEOUT window 410 step 1 layer 1 aux 600 | seq 1 flag 1 A 1 B 1
RELEASE window 410 ... | seq 1 flag 1 A 1 B 1
RELEASE-NOT-DRAINED window 410 ... aux 5000 | seq 1 flag 4294967295 A 4294967295 B 4294967295
The same in every hang I traced (5 runs). The CPU pool is not starved: it serves the layer in milliseconds.
The A/B settings suggested for this message in #649
Each was confirmed in the engine's own startup line or log, one per run, stopping at the first hang:
| setting | result |
|---|---|
STRATA_VERIFY_COHERENT=1 ("coherent words explicit") |
hung on run 1 |
STRATA_DOORBELL_STORE=1 ("doorbell stored") |
hung on run 1 |
STRATA_RESIDENT_PIN=0 |
does not apply here (only with --resident-experts / --resident-budget-gib); instead STRATA_ARENA_PIN_GIB=1 ("0 slices pinned (0 GiB)", the whole arena pageable): hung on run 1 |
What the GPU is doing at the hang
- Busy:
gpu_busy_percentof the primary card (0000:0f:00.0) is 98-99% for the whole hang; the helper card is at 0%. - Kernel log: no amdgpu / kfd / ttm messages at all.
- rocgdb: I attached with rocgdb (this needed an opt-in
prctl(PR_SET_PTRACER)in a diagnostic build). One HSA queue of the primary card has 2,781 packets pending. The packet at its read index is a dispatch ofgr_norm_split_kernel, grid [1536,4,1], which is the hyper-connection read for this window's 6 tokens. Another HSA queue has 376 packets pending, and its head packet is not a dispatch. The process has no AMDGPU waves at all. Together with the 99% busy reading, the command processor is waiting on something; no shader is spinning. - Attach changes the outcome: in the two runs where rocgdb attached, the GPU finished about 2.9 s after the release (
RELEASE-DRAINED). In the four runs without it, it did not finish within 5 s. Attaching makes KFD unmap and remap the process's queues, so I could not read the head packets' signals with rocgdb before the state changed.
Bisection
One setting per arm, up to 4 runs per arm, stopping at the first hang:
| arm | hangs |
|---|---|
| stock, default settings | 5 of 5 |
--adapt-every 100000 |
0 of 7 |
primary card's adaptive swaps only (helper's RemoteExpertOpt::adapt switched off in a diagnostic build) |
hung on run 1 |
helper's RemoteExpertOpt::adapt only (primary's swaps switched off) |
0 of 4 |
HSA_ENABLE_SDMA=0 |
0 of 4 |
STRATA_SH_STREAM=0 |
hung on run 2 |
STRATA_PREFILL_MMQ=0 (same -DSTRATA_PREFILL_MMQ=ON binary; checked in the engine's /proc/<pid>/environ) |
0 of 4 |
During the hangs, VRAM and GTT use were flat (sampled every ~1.6 s), and amd-evicted-vram in the engine's DRM fdinfo stayed 0.
Reading (not proven)
Three things are each necessary: the primary card's adaptive swaps (cudaMemcpyAsync H2D on adapt_stream, run on an SDMA engine), SDMA itself, and the MMQ prompt path. Both prompt paths stream the expert blobs over the copy stream, but the MMQ path adds per-group cross-stream waits and events (#372). The sequence is: swaps land in cache slots during the previous request's decode; the next prompt borrows cache slots and MMQ runs its groups there; right after the slots are refilled, the first verify window stalls. At the hang, the primary card's command processor sits on a cross-queue wait that queue remapping releases.
What I have not shown is whether the signal it waits on has already completed (a stale read or lost wakeup on the ROCm side) or genuinely has not (the copy, or an event dependency). Next I plan to read the head packets and their signal values from inside the process through an HSA_TOOLS_LIB shim, because rocgdb disturbs the state. Also not explained: why it needs about 7 earlier requests.
Cost of the three settings that avoid it
Same requests, means of 3-4 runs:
--adapt-every 100000: the prompt read is unchanged; decode is 3-11% slower (54-62 vs 60-69 t/s).HSA_ENABLE_SDMA=0: decode is unchanged; the prompt read is 10-24% slower (short prompts 219-235 vs 287-308 t/s, 4K 385-405 vs 450-459, 31K 426 vs 471; the 31K baseline is the--adapt-every 100000arm's, since the stock runs hung there, and that switch does not touch the prompt read).STRATA_PREFILL_MMQ=0: decode is unchanged; the prompt read is 28-42% slower (short prompts 175-179 vs 285-307 t/s, 4K 297-314 vs 449-458, 31K 339 vs 471).
Not #649
#649 prints the same message, but there it hangs in decode at varying layers, on --resident-budget-gib with 31 GB of RAM, and STRATA_RESIDENT_PIN=0 avoided it. Here it hangs right after a long prompt, always between layers 0 and 1, with 128 GB of RAM, and an unpinned arena does not avoid it.
I can run an instrumented build here, or share the full logs and the request script.
Edited 2026-10-05: the stock run count was corrected from "8 of 10" to 5 of 5 at default settings (the earlier figure mixed in runs with other settings); the prompt-read hang is attributed to the STRATA_SH_STREAM=0 arm, and the 31K baseline of the cost section is marked as the --adapt-every arm's.
- Dominant language
- C++
- Stars
- 11.6k
- Forks
- 1k
- Avg merge
- 7h 46m
- Merged PRs (30d)
- 30
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Niko1221/Strata
-
Difficulty 1/5 Under an hour Newbie friendliness 82/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 80/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 64/100
Niko1221/Strata#1248 · 1 comment ·
Maintainers usually reply within 1 day
Similar issues
-
8-membered-ring atrop stereo lost in 2026.09.1Possibly taken A pull request linked to this issue is open or already merged. Openbug
Difficulty 2/5 Half a day Newbie friendliness 86/100
Maintainers usually reply within 2 days
-
thread safetyOpen1.0
Difficulty 1/5 Under an hour Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
libasr headers?Open
Difficulty 2/5 1-3 hours Newbie friendliness 66/100
Maintainers usually reply within 1 day
-
Type: Bug :bug:
Difficulty 2/5 1-3 hours Newbie friendliness 83/100
-
llvm-trunk
Difficulty 1/5 Under an hour Newbie friendliness 88/100
Maintainers usually reply within 1 day