MI50 32 GB (gfx906) on 0.1.40.1: 126K/252K needles, a 16 GB-limit run, temperatures (results)
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 1/5
- Estimated time
- Under an hour
- Newbie friendliness
- 72/100
- Issue type
- Documentation
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- cpp
- Domain
- performance
Research direction
This is a benchmark report, with no code change or test named. Read the existing results in bench/results/2026-10-04-community-mi50 and compare the inline hardware, configuration, and measurements. Done means deciding whether to add the raw logs and scripts as a bench/results folder; they are available from the reporter on request.
Written by the indexing model from the issue text.
Description
Results of a test of Strata 0.1.40.1 (82f46a8) on one AMD Instinct MI50 32 GB (gfx906), measured on 2026-10-06. This complements the existing MI50 result (bench/results/2026-10-04-community-mi50, 0.1.38, another host, --spec 3): a later engine version, an AVX-512 host, and it adds 252K, a run with the card limited to 16 GB, and temperature readings. Runs were driven by Claude Code on my own machine. I am posting it as an issue with the numbers and configuration inline (no files attached); the raw logs and scripts are available if you want them as a bench/results folder.
Short version: decode 37-55 tok/s on short prompts, 44.4 / 41.0 tok/s at 34K / 68K of context, prefill 409-428 tok/s from 8K to 119K tokens. With the full 32 GB the three-needle check passed at 126K (3/3) and 252K (3/3), decode 42.5 tok/s at 126K and 36.3 tok/s at 252K. With the card limited to 16 GB it passed at 126K (31.2 tok/s) but failed at 252K: the answer was !!!!... (0/3). That one was not reproduced (see "The 252K / 16 GB failure"). One run per case, synthetic text: these numbers do not establish answer quality or performance on other workloads.
Hardware and software
- MI50 32 GB (Vega 20, gfx906), 34,342,961,152 bytes of VRAM; the engine's PCIe probe: 25.1 GB/s host to device. No power cap set by me, clocks not changed.
- Intel Core i9-11900H (the OS reports it as "Genuine Intel(R) CPU 0000 @ 2.60GHz", up to 4.8 GHz), 8 cores / 16 threads, AVX-512; 62 GB RAM; Ubuntu 26.04.1 LTS, kernel 7.0.0-38-generic. 7 expert-pool workers on logical processors 1-7, host thread on 0.
- ROCm 7.2.2 (AMD clang 22.0.0git,
roc-7.2.2);lldneeded a compatiblelibxml2onLD_LIBRARY_PATHon this OS. - Strata
v0.1.40.1, commit82f46a8c8f475f001ad76d92f58f4a4f8ffb0253, built with-DSTRATA_HIP_GFX906=ONas indocs/AMD_HIP.md; the engine reports version 0.1.40. - Nothing else used the GPU; no system setting (power, driver, kernel) was changed.
Model and configuration
ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF, IQ2_XS (the GGUF already on the machine; pack and MTP layer built by the setup flow). Vision off, temperature 0, no reasoning. Engine arguments (paths removed), 32 GB / 131,072:
--pack <data-dir>/packs/iq2_xs --native <models-dir>/.../IQ2_XS/...-00001-of-00002.gguf --ple-gguf <models-dir>/.../...-00002-of-00002.gguf
--expert-profile <strata-dir>/data/expert-profile.bin --expert-cache auto --prefill auto
--spec 4 --spec-min-p 0.5 --mtp <data-dir>/mtp/rt --max-context 131072 --kv int8 --kv-resident 32768
The 262,144 runs use --max-context 262144; the "16 GB" runs add --vram-reserve-mib 16384 (= 32,752 - 16,368, what a 16 GB MI50 exposes), so they are a simulation by Strata's reserve, not a real 16 GB card. --kv-resident 32768 (KV streaming) is what the installer chose; I did not compare against it being off.
- 32 GB: 19,473 of 24,576 experts (~79%, 26.17 GiB) in VRAM, pre-filled from the profile, no eviction; decode cache hit rate 95.8-98.5%; ~39 GB of RAM in use; loading the 33.02 GiB of experts took ~1 min (3.68 GiB/s).
- 16 GB: 8,099 slots (~33%, 10.85 GiB); hit rate 73-91%.
Results (single runs; the engine's own strata serve: timings)
| Case | 32 GB | 16 GB (simulated) |
|---|---|---|
| Decode, short prompt (4 requests) | 37.2 / 54.2 / 55.1 / 40.6 tok/s (MTP drafts accepted 47-76%) | 30-34 tok/s |
| Decode at 34K context (350 tokens) | 44.4 tok/s | 35.7 tok/s |
| Decode at 68K context (350 tokens) | 41.0 tok/s | 38.9 tok/s |
| Prefill, 26-55-token prompts | 63-79 tok/s | 43-44 tok/s |
| Prefill, 8K to 119K tokens | 409-428 tok/s (client-side; 410-430 server-side) | 394-419 tok/s |
| Needle in the middle, 2K to 119K tokens | 6 of 6 | 6 of 6 |
Three 6-digit needles at 10 / 50 / 90 % depth, prompt near each limit; the second request re-sends the same text with another question and reuses the first one's prefix (114,688 tokens at 126K, 245,760 at 252K), so the last column is decode, not prefill:
| Mode | Context | Prompt tokens | Prefill | Needles | Decode with the full context |
|---|---|---|---|---|---|
| 32 GB | 131,072 | 125,991 | 409 tok/s (308 s) | 3 of 3 | 42.5 tok/s (hit 98.3%) |
| 32 GB | 262,144 | 252,304 | 370 tok/s (681 s) | 3 of 3 | 36.3 tok/s (hit 98.5%) |
| 16 GB | 131,072 | 125,991 | 403 tok/s (313 s) | 3 of 3 | 31.2 tok/s (hit 91.2%) |
| 16 GB | 262,144 | 252,304 | 382 tok/s (660 s) | 0 of 3: !!!!... |
256 tokens at 26.8 tok/s, no draft accepted |
Engine request lines for the four long runs
32 GB, 126K : prompt 125991 tokens = 0 reused + 125991 read in 307796 ms (409.3 tok/s), 23 generated in 524 ms (43.9 tok/s), drafts accepted 15 of 20
prompt 125985 tokens = 114688 reused + 11297 read in 31579 ms (357.7 tok/s), 293 generated in 6894 ms (42.5 tok/s), drafts accepted 144 of 236
32 GB, 252K : prompt 252304 tokens = 0 reused + 252304 read in 681378 ms (370.3 tok/s), 23 generated in 6606 ms (3.5 tok/s), drafts accepted 16 of 18
prompt 252298 tokens = 245760 reused + 6538 read in 21504 ms (304.0 tok/s), 300 generated in 8265 ms (36.3 tok/s), drafts accepted 141 of 255
16 GB, 126K : prompt 125991 tokens = 0 reused + 125991 read in 312521 ms (403.1 tok/s), 23 generated in 7644 ms (3.0 tok/s), drafts accepted 15 of 20
prompt 125985 tokens = 114688 reused + 11297 read in 31823 ms (355.0 tok/s), 300 generated in 9618 ms (31.2 tok/s), drafts accepted 134 of 252
16 GB, 252K : prompt 252304 tokens = 0 reused + 252304 read in 660486 ms (382.0 tok/s), 60 generated in 3576 ms (16.8 tok/s), drafts accepted 3 of 3
prompt 252298 tokens = 245760 reused + 6538 read in 20185 ms (323.9 tok/s), 256 generated in 9563 ms (26.8 tok/s), drafts accepted 0 of 0
Temperatures
Junction from rocm-smi every 10 s; memory temperature was not recorded. Prefill of 8K-119K reached 99-101 C junction at 155-200 W (edge 78 C) in the 32 GB mode and 89-97 C in the 16 GB mode; the 126K runs stayed at 89-97 C and the 252K runs reached 97-104 C (the 32 GB 252K run peaked at 103 C and passed). Back to 40-53 C at idle. My guard (stop at 104 C twice in a row) fired once, during a repeat of the 16 GB 252K run (the request ended with HTTP 503). The card's sysfs lists critical limits of 105 C (junction/edge) and 94 C (HBM).
The 252K / 16 GB failure
Output !!!!... and 0 of 3 needles at 252,304 tokens; the server logged no error and dmesg showed no GPU reset. The same prompt passed in the full 32 GB mode. The repeat was cut by my temperature guard, so I do not know whether it reproduces. The same !!!! pattern at long context is reported on other cards (#606, #879, #871); I found no earlier gfx906 report. I have a hypothesis, not verified, that heat in the HBM was involved (junction reached 102-104 C in both 16 GB 252K runs); there is no direct evidence, and a different expert cache (8,099 slots) may matter too.
Notes for other MI50 users
- On ROCm 7.2.2, v0.1.40.1 did not build with
-DSTRATA_HIP_GFX906=ONuntil three small fixes: a barereturn;in aboolfunction infused_gr.cu, and two#ifguards inmtp.cppandvmm.cpp. I looked atmain(d5ea71337): all three are already fixed there. I have not rebuilt or re-runmain. setup.pyhas nogfx906inAMD_ARCHS(checked onmain), so the installer does not recognise the card: I wroteengine/BUILD.jsonby hand (source: local-hip-gfx906,archs: [gfx906], the source hash) so setup accepted the engine.- The full card held 19,473 of 24,576 experts at both 131,072 and 262,144; the 252K prefill ran at 370 tok/s, about 10% below the 126K one.
Limits
One quant (IQ2_XS), one run per case, synthetic repetitive text (simple needles), no --parallel, no soak, no power limit, no memory-temperature log. The 16 GB rows come from a reserve, not a real 16 GB card. The numbers come from the engine version and host above.
- Dominant language
- C++
- Stars
- 11.6k
- Forks
- 1k
- Avg merge
- 7h 46m
- Merged PRs (30d)
- 30
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Niko1221/Strata
-
Difficulty 2/5 1-3 hours Newbie friendliness 66/100
Maintainers usually reply within 1 day
-
expert_cache_segmented_test fails on HIP builds instead of skipping (--vram-elastic is CUDA-only)Possibly taken @nekomario28 claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 83/100
Maintainers usually reply within 1 day
-
hip_q2_zero fails on gfx1201 (R9700) with ROCm 7.10: Q2_0 signed-zero fix e9a5f8d is gated to gfx1012 / HIP < 7Possibly taken @nekomario28 claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 66/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Niko1221/Strata#1322 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 82/100
Maintainers usually reply within 1 day
Similar issues
-
agent:Windows bug
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
? - Needs Triage bot_watch bug
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
NVIDIA/cudf-spark-jni#5267 · 1 comment ·
Maintainers usually reply within 1 day
-
area/docdb kind/bug priority/medium status/awaiting-triage
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
yugabyte/yugabyte-db#34873 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day