Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

MI50 32 GB (gfx906) on 0.1.40.1: 126K/252K needles, a 16 GB-limit run, temperatures (results)

Open Beginner friendly
#1,463 2 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
1/5
Estimated time
Under an hour
Newbie friendliness
72/100
Issue type
Documentation
Clarity
Clearly specified
Activity status
Active
Tech stack
cpp
Domain
performance

Research direction

This is a benchmark report, with no code change or test named. Read the existing results in bench/results/2026-10-04-community-mi50 and compare the inline hardware, configuration, and measurements. Done means deciding whether to add the raw logs and scripts as a bench/results folder; they are available from the reporter on request.

Written by the indexing model from the issue text.

Description

Results of a test of Strata 0.1.40.1 (82f46a8) on one AMD Instinct MI50 32 GB (gfx906), measured on 2026-10-06. This complements the existing MI50 result (bench/results/2026-10-04-community-mi50, 0.1.38, another host, --spec 3): a later engine version, an AVX-512 host, and it adds 252K, a run with the card limited to 16 GB, and temperature readings. Runs were driven by Claude Code on my own machine. I am posting it as an issue with the numbers and configuration inline (no files attached); the raw logs and scripts are available if you want them as a bench/results folder.

Short version: decode 37-55 tok/s on short prompts, 44.4 / 41.0 tok/s at 34K / 68K of context, prefill 409-428 tok/s from 8K to 119K tokens. With the full 32 GB the three-needle check passed at 126K (3/3) and 252K (3/3), decode 42.5 tok/s at 126K and 36.3 tok/s at 252K. With the card limited to 16 GB it passed at 126K (31.2 tok/s) but failed at 252K: the answer was !!!!... (0/3). That one was not reproduced (see "The 252K / 16 GB failure"). One run per case, synthetic text: these numbers do not establish answer quality or performance on other workloads.

Hardware and software

  • MI50 32 GB (Vega 20, gfx906), 34,342,961,152 bytes of VRAM; the engine's PCIe probe: 25.1 GB/s host to device. No power cap set by me, clocks not changed.
  • Intel Core i9-11900H (the OS reports it as "Genuine Intel(R) CPU 0000 @ 2.60GHz", up to 4.8 GHz), 8 cores / 16 threads, AVX-512; 62 GB RAM; Ubuntu 26.04.1 LTS, kernel 7.0.0-38-generic. 7 expert-pool workers on logical processors 1-7, host thread on 0.
  • ROCm 7.2.2 (AMD clang 22.0.0git, roc-7.2.2); lld needed a compatible libxml2 on LD_LIBRARY_PATH on this OS.
  • Strata v0.1.40.1, commit 82f46a8c8f475f001ad76d92f58f4a4f8ffb0253, built with -DSTRATA_HIP_GFX906=ON as in docs/AMD_HIP.md; the engine reports version 0.1.40.
  • Nothing else used the GPU; no system setting (power, driver, kernel) was changed.

Model and configuration

ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF, IQ2_XS (the GGUF already on the machine; pack and MTP layer built by the setup flow). Vision off, temperature 0, no reasoning. Engine arguments (paths removed), 32 GB / 131,072:

--pack <data-dir>/packs/iq2_xs --native <models-dir>/.../IQ2_XS/...-00001-of-00002.gguf --ple-gguf <models-dir>/.../...-00002-of-00002.gguf
--expert-profile <strata-dir>/data/expert-profile.bin --expert-cache auto --prefill auto
--spec 4 --spec-min-p 0.5 --mtp <data-dir>/mtp/rt --max-context 131072 --kv int8 --kv-resident 32768

The 262,144 runs use --max-context 262144; the "16 GB" runs add --vram-reserve-mib 16384 (= 32,752 - 16,368, what a 16 GB MI50 exposes), so they are a simulation by Strata's reserve, not a real 16 GB card. --kv-resident 32768 (KV streaming) is what the installer chose; I did not compare against it being off.

  • 32 GB: 19,473 of 24,576 experts (~79%, 26.17 GiB) in VRAM, pre-filled from the profile, no eviction; decode cache hit rate 95.8-98.5%; ~39 GB of RAM in use; loading the 33.02 GiB of experts took ~1 min (3.68 GiB/s).
  • 16 GB: 8,099 slots (~33%, 10.85 GiB); hit rate 73-91%.

Results (single runs; the engine's own strata serve: timings)

Case 32 GB 16 GB (simulated)
Decode, short prompt (4 requests) 37.2 / 54.2 / 55.1 / 40.6 tok/s (MTP drafts accepted 47-76%) 30-34 tok/s
Decode at 34K context (350 tokens) 44.4 tok/s 35.7 tok/s
Decode at 68K context (350 tokens) 41.0 tok/s 38.9 tok/s
Prefill, 26-55-token prompts 63-79 tok/s 43-44 tok/s
Prefill, 8K to 119K tokens 409-428 tok/s (client-side; 410-430 server-side) 394-419 tok/s
Needle in the middle, 2K to 119K tokens 6 of 6 6 of 6

Three 6-digit needles at 10 / 50 / 90 % depth, prompt near each limit; the second request re-sends the same text with another question and reuses the first one's prefix (114,688 tokens at 126K, 245,760 at 252K), so the last column is decode, not prefill:

Mode Context Prompt tokens Prefill Needles Decode with the full context
32 GB 131,072 125,991 409 tok/s (308 s) 3 of 3 42.5 tok/s (hit 98.3%)
32 GB 262,144 252,304 370 tok/s (681 s) 3 of 3 36.3 tok/s (hit 98.5%)
16 GB 131,072 125,991 403 tok/s (313 s) 3 of 3 31.2 tok/s (hit 91.2%)
16 GB 262,144 252,304 382 tok/s (660 s) 0 of 3: !!!!... 256 tokens at 26.8 tok/s, no draft accepted
Engine request lines for the four long runs
32 GB, 126K : prompt 125991 tokens = 0 reused + 125991 read in 307796 ms (409.3 tok/s), 23 generated in 524 ms (43.9 tok/s), drafts accepted 15 of 20
              prompt 125985 tokens = 114688 reused + 11297 read in 31579 ms (357.7 tok/s), 293 generated in 6894 ms (42.5 tok/s), drafts accepted 144 of 236
32 GB, 252K : prompt 252304 tokens = 0 reused + 252304 read in 681378 ms (370.3 tok/s), 23 generated in 6606 ms (3.5 tok/s), drafts accepted 16 of 18
              prompt 252298 tokens = 245760 reused + 6538 read in 21504 ms (304.0 tok/s), 300 generated in 8265 ms (36.3 tok/s), drafts accepted 141 of 255
16 GB, 126K : prompt 125991 tokens = 0 reused + 125991 read in 312521 ms (403.1 tok/s), 23 generated in 7644 ms (3.0 tok/s), drafts accepted 15 of 20
              prompt 125985 tokens = 114688 reused + 11297 read in 31823 ms (355.0 tok/s), 300 generated in 9618 ms (31.2 tok/s), drafts accepted 134 of 252
16 GB, 252K : prompt 252304 tokens = 0 reused + 252304 read in 660486 ms (382.0 tok/s), 60 generated in 3576 ms (16.8 tok/s), drafts accepted 3 of 3
              prompt 252298 tokens = 245760 reused + 6538 read in 20185 ms (323.9 tok/s), 256 generated in 9563 ms (26.8 tok/s), drafts accepted 0 of 0

Temperatures

Junction from rocm-smi every 10 s; memory temperature was not recorded. Prefill of 8K-119K reached 99-101 C junction at 155-200 W (edge 78 C) in the 32 GB mode and 89-97 C in the 16 GB mode; the 126K runs stayed at 89-97 C and the 252K runs reached 97-104 C (the 32 GB 252K run peaked at 103 C and passed). Back to 40-53 C at idle. My guard (stop at 104 C twice in a row) fired once, during a repeat of the 16 GB 252K run (the request ended with HTTP 503). The card's sysfs lists critical limits of 105 C (junction/edge) and 94 C (HBM).

The 252K / 16 GB failure

Output !!!!... and 0 of 3 needles at 252,304 tokens; the server logged no error and dmesg showed no GPU reset. The same prompt passed in the full 32 GB mode. The repeat was cut by my temperature guard, so I do not know whether it reproduces. The same !!!! pattern at long context is reported on other cards (#606, #879, #871); I found no earlier gfx906 report. I have a hypothesis, not verified, that heat in the HBM was involved (junction reached 102-104 C in both 16 GB 252K runs); there is no direct evidence, and a different expert cache (8,099 slots) may matter too.

Notes for other MI50 users

  • On ROCm 7.2.2, v0.1.40.1 did not build with -DSTRATA_HIP_GFX906=ON until three small fixes: a bare return; in a bool function in fused_gr.cu, and two #if guards in mtp.cpp and vmm.cpp. I looked at main (d5ea71337): all three are already fixed there. I have not rebuilt or re-run main.
  • setup.py has no gfx906 in AMD_ARCHS (checked on main), so the installer does not recognise the card: I wrote engine/BUILD.json by hand (source: local-hip-gfx906, archs: [gfx906], the source hash) so setup accepted the engine.
  • The full card held 19,473 of 24,576 experts at both 131,072 and 262,144; the 252K prefill ran at 370 tok/s, about 10% below the 126K one.

Limits

One quant (IQ2_XS), one run per case, synthetic repetitive text (simple needles), no --parallel, no soak, no power limit, no memory-temperature log. The 16 GB rows come from a reserve, not a real 16 GB card. The numbers come from the engine version and host above.

Dominant language
C++
Stars
11.6k
Forks
1k
Avg merge
7h 46m
Merged PRs (30d)
30

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Niko1221/Strata

All issues in Niko1221/Strata

Similar issues

More C++ issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.