bench: community report, 2x Tesla V100-PCIE-32GB (sm_70, CUDA 12 source build): IQ3_S at 512K with --layer-split, and what the new decode profiler says about where tok/s comes from
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 62/100
- Issue type
- Documentation
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- cpp
- Domain
- documentation
Research direction
This is a benchmark report with two explicitly suggested doc changes: extend the --lookup-chain CLI help to say it targets templated/repetitive content (the reporter measured a net loss on prose/code), and clarify --pool-affinity options in relation to NUMA on multi-socket boxes. Start from where the --lookup-chain flag's help string is defined in the argument-parsing code, then check docs/BATCHING.md and docs/ for the RAM-speed and layer-split notes the report references. Done when the help text states the intended workload; the linked external write-up and the decode profiler tables are context, not required edits.
Written by the indexing model from the issue text.
Description
What this is: a two-card V100-PCIE-32GB Linux source build (sm_70, CUDA 12) running Qwen3.8-Flash-Next IQ3_S at the full 524,288-token context with vision on. The 0.1.40 notes asked for Volta results; the existing Volta community reports are 16 GB cards or mixed Volta+Pascal pairs, so a 32 GB pair at 512K may be a useful data point.
The headline is not a speed number — it is that decode tok/s on this engine is governed by tokens per verify window, which is a property of the content, not of the machine, and that the engine can now show you this directly with the profiler added in 0.1.40.
Hardware
- 2× Tesla V100-PCIE-32GB (sm_70), passive, server-cooled. GPU0 ~29.5 GiB used, GPU1 ~32.2 GiB used (of 32,768 MiB each).
nvidia-smi topo -m: GPU0↔GPU1 = SYS — different CPU sockets, no P2P, no NVLink. Layer split (--layer-split) is the working path, as the docs describe.- CPU: 2× Xeon E5-2673 v3 (12 cores each, 48 threads total, AVX2 only — no AVX-512). PCIe Gen3 x16.
- Memory: 125 GiB, DDR3-1600 (see the note at the end — this is a real shortcoming of this box and it turned out not to matter).
- Linux, self-built engine with
-DCMAKE_CUDA_ARCHITECTURES=70 -DSTRATA_EXPERIMENTAL_SM60=ON, CUDA 12.x.
Configuration
--max-context 524288 --rope-scaling yarn --rope-scale 2 --kv int8 --kv-resident 20480 --spec 4 --spec-min-p 0.70 --ple-io mmap --layer-split 20 --vision --vram-reserve-mib 700
Two notes on the reasoning, both measured rather than assumed:
--ple-io mmap, notdirectorram.ramis structurally wrong here, not just slower: locking the 26.8 GiB PLE table drivesMemFreeto 603 MB and the kernel then reclaims the process's own anonymous pages (VmSwap1.45 GiB,pswpin+59%,compact_stall324,567). The symptom is nasty because it does not look like a performance problem — the engine's own decode rate stays normal while a client sees tens of seconds of wall clock.directpays an SSD read on every token's table lookup.mmapmaps it into the page cache without locking, so rows in use stay resident and the rest stays reclaimable (VmLck 0,VmSwap 0). We are not quoting a percentage for mmap over direct — the arms were measured in separate windows and cannot be separated from run-to-run spread.- 512K is configured and works, with
--kv-residentas a hard requirement rather than an optimisation. Without it the KV state takes the expert slots.
Decode: the profiler explains what moves the number
Setting STRATA_DECODE_TIMING=1 (with STRATA_VERIFY_PROFILE=1) logs a per-window breakdown per request. This is a genuinely useful addition — it is what let us stop guessing:
strata decode timing: 281 windows, avg T 2.04, 1.82 tokens/window, 25.80 ms/window
= verify 23.56 (GPU-reach wait 0.00 + per-layer host 0.00 [plan 0.05 actq 0.07 jobs 0.00 CPU 0.51]
+ stage 0.06) + commit/emit 0.37 + draft 1.59
per layer-window: CPU experts 0.10 (0.13 entries), VRAM hits 11.74, PCIe 0.01
Across seven profiled runs:
| tokens/window | ms/window | ms/token | CPU experts | VRAM hits |
|---|---|---|---|---|
| 3.09 | 37.89 | 12.26 | 0.50 | 18.99 |
| 2.79 | 33.49 | 12.00 | 0.41 | 17.08 |
| 2.55 | 33.19 | 13.02 | 0.35 | 16.46 |
| 1.84 | 25.80 | 14.02 | 0.20 | 12.18 |
| 1.59 | 24.06 | 15.13 | 0.22 | 9.86 |
| 1.50 | 23.58 | 15.72 | 0.12 | 9.41 |
ms/token is flat (12.0–15.7 ms) and ms/window is sub-linear in tokens/window — about 23 ms fixed per window plus ~9 ms for each token the verifier accepts. So throughput is set by tokens per window, which is how many draft tokens the verifier accepts on this content. The CPU expert pool is 0.5–4% of window time and PCIe is 0.01 ms: on this box neither is close to being the bottleneck, even at 512K with a DDR3-1600 CPU half.
Decode by content type — interleaved
Because a single number is not reproducible across content, here is the number with its content attached. 12 rounds per type, three types interleaved within one service lifetime (A,B,C,A,B,C…), first 2 rounds discarded, medians reported, and the whole thing run twice on separate occasions:
| Content | Round 1 (median) | Round 2 (median) | Acceptance |
|---|---|---|---|
| prose (zh) | 84.8 | 80.8 | 78.8% |
| code | 82.8 | 79.3 | 83.0% |
| prose (en) | 70.9 | 68.3 | 81.5% |
- Content-to-content spread is 1.2×; the same content across the two rounds differs by ≤5%, and the ordering reproduced.
- Acceptance rate does not predict tok/s — the lowest-acceptance arm (Chinese prose, 78.8%) was the fastest. What matters is tokens/window, i.e. acceptance combined with draft depth, not acceptance alone.
- The ordering is not the intuitive one (we expected code to win). Worth measuring per workload rather than assuming.
Prefill (cold, engine counters)
| Prompt | Tokens read | Engine prompt time | Throughput |
|---|---|---|---|
| ~256 | 315 | 1,736 ms | 181.5 tok/s |
| ~1K | 1,004 | 1,892 ms | 530.8 tok/s |
| ~4K | 3,707 | 3,425 ms | 1,082.3 tok/s |
| ~16K | 14,572 | 8,653 ms | 1,684.1 tok/s |
Fresh prompt per row — a repeat is served from the conversation cache and reports a meaningless figure. Above 16K we could not measure cold without changing the configuration, so we do not claim those rows.
Other things that may be worth passing upstream
--spec 4beat--spec 8on this box (79.4 vs 72.9 tok/s at 256K, acceptance 76–78% vs 61–65%) — shallow drafts, higher hit rate. The--spec-min-p 0.70row matched your calibration guidance.--batchwas a net loss here and the reason is indocs/BATCHING.md: batch windows carry no MTP drafts, and this configuration depends on MTP heavily. Four concurrent requests queued at 0.90× — consistent with "one request at a time".--lookup-chain(opt-in, default 0): we measured this as a net loss on prose and code prompts (73.7 vs 77.0). Reading the help, it drafts by matching repeated context, so it should only pay on templated/repetitive content — which neither of our benchmark prompts is. Not a bug report, just a note that "no measurable gain" here is expected rather than surprising, and it might be worth a line in the help saying which workloads it targets.--pool-affinityhas no "restrict to one NUMA node" mode (onlyall/auto/p-cores). Wrapping the service innumactl --cpunodebind=0was worse (49–69 tok/s), so we left the defaults alone. Noting it because "stop the workers crossing QPI" looks like an obvious win on a two-socket box and there is no supported way to try it.
What we did not measure
- Full 512K recall. Our deepest retrieval test is 241,255 tokens (correct). 512K is configured and answers, but recall at the very top of the window is untested.
- Quality benchmarks. No GSM8K/HumanEval/MMLU. Decode and prefill numbers here are safe to quote; quality numbers are absent.
- A decode cost for 512K. We measured 256K and 512K and the difference sits inside the run-to-run spread, so we make no claim either way.
- Long-run stability. These are figures from a live server, not a multi-day soak.
A note on the DDR3-1600 finding
This box runs DDR3-1600 although the E5-2673 v3 supports DDR4-2133, and docs/ warns that RAM below its rated speed slows the CPU half. It is a real shortcoming, and the profiler is what showed it does not matter here: the CPU expert pool is 0.5–4% of window time, so even doubling memory bandwidth has a few percent of headroom. Recording it because "check your RAM speed" is the intuitive first move on a box like this and the profiler says it is not where the time goes.
Full configuration, build notes and the complete negative-results log (including several of our own earlier claims that we had to withdraw) are in our write-up: https://github.com/ZackO2o/Strata-Qwen3.8-Flash-Next-512k-v100x2
Happy to re-run anything on request — the box is up and the profiler is one environment variable away.
- Dominant language
- C++
- Stars
- 11.6k
- Forks
- 1k
- Avg merge
- 7h 46m
- Merged PRs (30d)
- 30
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Niko1221/Strata
-
Difficulty 1/5 Under an hour Newbie friendliness 82/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 80/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 64/100
Niko1221/Strata#1248 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 1 day
Similar issues
-
bug iOS 🍎 ui/ux
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
MerginMaps/mobile#4744 ·
Maintainers usually reply within 1 day
-
Component: Ruby Type: bug
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 88/100
Maintainers usually reply within 3 days
-
Difficulty 1/5 Under an hour Newbie friendliness 92/100
kokkos/kokkos-kernels#3328 ·
Maintainers usually reply within 1 day
-
bug needs triage tcp
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
project-chip/connectedhomeip#74644 ·
Maintainers usually reply within 1 day