Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

HIP gfx1030, 2x RX 6900 XT: verify hangs after a long prompt when MMQ prompt path, adaptive expert swaps and SDMA meet (any one off avoids it)

Open
#884 3 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
25/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
cpp, linux

Research direction

Start by reproducing the hang with the request sequence and settings in the issue, then read the HSA queue head packets and signal values from inside the process through an HSA_TOOLS_LIB shim. Compare those signals with the reported cross-queue wait to determine whether it has completed or is still waiting on a dependency. Done means identifying the cause of the hang; the issue does not name specific source files or tests.

Written by the indexing model from the issue text.

Description

Summary

On 2x RX 6900 XT (gfx1030, HIP, Linux), engine 0.1.39 hangs on the first verify window after a ~31K-token prompt, once the engine has served about 7 earlier requests: verify: timed out at layer 1; its GPU waits were released but the GPU did not finish within 5 s (#267). Stock 6f32ec0 at default settings hung in 5 of 5 runs of the sequence below (and in 11 of 12 more runs of it with one diagnostic setting that did not avoid it, listed further down).

It needs three things at once; taking away any one of them stops it in my runs: the MMQ prompt path (STRATA_PREFILL_MMQ=0: 0/4), the primary card's adaptive expert swaps (--adapt-every 100000: 0/7), and SDMA copies (HSA_ENABLE_SDMA=0: 0/4). At the hang the primary card is 98-99% busy with no waves resident: its command processor is stuck at a wait, not a kernel. The root cause is not known yet; what is measured and what is not is listed below.

This is not #649, although the message is the same; see the last section.

Environment

  • Strata 6f32ec0 (engine 0.1.39), build-hip with -DSTRATA_PREFILL_MMQ=ON (as setup.py builds for AMD), gfx1030
  • 2x RX 6900 XT 16 GB, PCIe 4.0 x8 each; Ryzen 5 5600X; 128 GB DDR4 (about 69 GB still free at the hang)
  • Linux 7.0, ROCm 10.0
  • Model: Qwen3.8-Flash-Next GSQ-RCO IQ3_S, native pack, --ple-gguf, --mtp draft head
  • Arguments: --expert-cache auto --prefill auto --spec 4 --spec-min-p 0.5 --max-context 131072 --kv int8 --kv-resident 32768 --expert-cache-device1 auto --remote-expert-opt --pcie-frac 0, HIP_VISIBLE_DEVICES=1,0 (the expert arena, 46.84 GiB, is loaded whole and cudaHostRegistered)

Reproduction

A fresh server, then 8 chat requests, one after the other (greedy, thinking off, max_tokens 256): 7 short ones of 1.0K-4.9K tokens, then one of 31,219 tokens. The texts are slices of the repository's own docs and sources; the content does not matter (other slices hang the same way).

  • Stock, default settings, this sequence: 5 of 5 runs hung on the 31K request. With --prompt-cache 12 instead of the default 6: 1 of 2.
  • With 5 or fewer earlier requests before the 31K one: 0 of 5 (one of those with --prompt-cache 2).
  • Every stock hang was the first verify window after the batched prompt (6 tokens at position 31212). The one hang in the batched prompt read itself - no progress for 60 s ... (reading the prompt (batched): waiting for the GPU (attention, router) at layer 1 of the prompt chunk from token 0) - was in the STRATA_SH_STREAM=0 arm below.

STRATA_VERIFY_TRACE=1

The host serves layer 0 within about 3.3 ms (flag 1 A 1 B 1). The GPU's breadcrumbs stop after layer 0's waitCPU; combine is never stamped:

layer 0 group 0: pre(0)=0.0 hc-read0(1)=101.6 ... waitB(21)=661.2 PCIe grp(22)=672.4 waitCPU(23)=3849.1
TIMEOUT  window 410 step 1 layer 1 aux 600 | seq 1 flag 1 A 1 B 1
RELEASE  window 410 ... | seq 1 flag 1 A 1 B 1
RELEASE-NOT-DRAINED  window 410 ... aux 5000 | seq 1 flag 4294967295 A 4294967295 B 4294967295

The same in every hang I traced (5 runs). The CPU pool is not starved: it serves the layer in milliseconds.

The A/B settings suggested for this message in #649

Each was confirmed in the engine's own startup line or log, one per run, stopping at the first hang:

setting result
STRATA_VERIFY_COHERENT=1 ("coherent words explicit") hung on run 1
STRATA_DOORBELL_STORE=1 ("doorbell stored") hung on run 1
STRATA_RESIDENT_PIN=0 does not apply here (only with --resident-experts / --resident-budget-gib); instead STRATA_ARENA_PIN_GIB=1 ("0 slices pinned (0 GiB)", the whole arena pageable): hung on run 1

What the GPU is doing at the hang

  • Busy: gpu_busy_percent of the primary card (0000:0f:00.0) is 98-99% for the whole hang; the helper card is at 0%.
  • Kernel log: no amdgpu / kfd / ttm messages at all.
  • rocgdb: I attached with rocgdb (this needed an opt-in prctl(PR_SET_PTRACER) in a diagnostic build). One HSA queue of the primary card has 2,781 packets pending. The packet at its read index is a dispatch of gr_norm_split_kernel, grid [1536,4,1], which is the hyper-connection read for this window's 6 tokens. Another HSA queue has 376 packets pending, and its head packet is not a dispatch. The process has no AMDGPU waves at all. Together with the 99% busy reading, the command processor is waiting on something; no shader is spinning.
  • Attach changes the outcome: in the two runs where rocgdb attached, the GPU finished about 2.9 s after the release (RELEASE-DRAINED). In the four runs without it, it did not finish within 5 s. Attaching makes KFD unmap and remap the process's queues, so I could not read the head packets' signals with rocgdb before the state changed.

Bisection

One setting per arm, up to 4 runs per arm, stopping at the first hang:

arm hangs
stock, default settings 5 of 5
--adapt-every 100000 0 of 7
primary card's adaptive swaps only (helper's RemoteExpertOpt::adapt switched off in a diagnostic build) hung on run 1
helper's RemoteExpertOpt::adapt only (primary's swaps switched off) 0 of 4
HSA_ENABLE_SDMA=0 0 of 4
STRATA_SH_STREAM=0 hung on run 2
STRATA_PREFILL_MMQ=0 (same -DSTRATA_PREFILL_MMQ=ON binary; checked in the engine's /proc/<pid>/environ) 0 of 4

During the hangs, VRAM and GTT use were flat (sampled every ~1.6 s), and amd-evicted-vram in the engine's DRM fdinfo stayed 0.

Reading (not proven)

Three things are each necessary: the primary card's adaptive swaps (cudaMemcpyAsync H2D on adapt_stream, run on an SDMA engine), SDMA itself, and the MMQ prompt path. Both prompt paths stream the expert blobs over the copy stream, but the MMQ path adds per-group cross-stream waits and events (#372). The sequence is: swaps land in cache slots during the previous request's decode; the next prompt borrows cache slots and MMQ runs its groups there; right after the slots are refilled, the first verify window stalls. At the hang, the primary card's command processor sits on a cross-queue wait that queue remapping releases.

What I have not shown is whether the signal it waits on has already completed (a stale read or lost wakeup on the ROCm side) or genuinely has not (the copy, or an event dependency). Next I plan to read the head packets and their signal values from inside the process through an HSA_TOOLS_LIB shim, because rocgdb disturbs the state. Also not explained: why it needs about 7 earlier requests.

Cost of the three settings that avoid it

Same requests, means of 3-4 runs:

  • --adapt-every 100000: the prompt read is unchanged; decode is 3-11% slower (54-62 vs 60-69 t/s).
  • HSA_ENABLE_SDMA=0: decode is unchanged; the prompt read is 10-24% slower (short prompts 219-235 vs 287-308 t/s, 4K 385-405 vs 450-459, 31K 426 vs 471; the 31K baseline is the --adapt-every 100000 arm's, since the stock runs hung there, and that switch does not touch the prompt read).
  • STRATA_PREFILL_MMQ=0: decode is unchanged; the prompt read is 28-42% slower (short prompts 175-179 vs 285-307 t/s, 4K 297-314 vs 449-458, 31K 339 vs 471).

Not #649

#649 prints the same message, but there it hangs in decode at varying layers, on --resident-budget-gib with 31 GB of RAM, and STRATA_RESIDENT_PIN=0 avoided it. Here it hangs right after a long prompt, always between layers 0 and 1, with 128 GB of RAM, and an unpinned arena does not avoid it.

I can run an instrumented build here, or share the full logs and the request script.

Edited 2026-10-05: the stock run count was corrected from "8 of 10" to 5 of 5 at default settings (the earlier figure mixed in runs with other settings); the prompt-read hang is attributed to the STRATA_SH_STREAM=0 arm, and the 31K baseline of the cost section is marked as the --adapt-every arm's.

Dominant language
C++
Stars
11.6k
Forks
1k
Avg merge
7h 46m
Merged PRs (30d)
30

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Niko1221/Strata

All issues in Niko1221/Strata

Similar issues

More C++ issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.