Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

bench: community report, 2x Tesla V100-PCIE-32GB (sm_70, CUDA 12 source build): IQ3_S at 512K with --layer-split, and what the new decode profiler says about where tok/s comes from

Open Beginner friendly
#1,300 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
62/100
Issue type
Documentation
Clarity
Mostly clear
Activity status
Active
Tech stack
cpp
Domain
documentation

Research direction

This is a benchmark report with two explicitly suggested doc changes: extend the --lookup-chain CLI help to say it targets templated/repetitive content (the reporter measured a net loss on prose/code), and clarify --pool-affinity options in relation to NUMA on multi-socket boxes. Start from where the --lookup-chain flag's help string is defined in the argument-parsing code, then check docs/BATCHING.md and docs/ for the RAM-speed and layer-split notes the report references. Done when the help text states the intended workload; the linked external write-up and the decode profiler tables are context, not required edits.

Written by the indexing model from the issue text.

Description

What this is: a two-card V100-PCIE-32GB Linux source build (sm_70, CUDA 12) running Qwen3.8-Flash-Next IQ3_S at the full 524,288-token context with vision on. The 0.1.40 notes asked for Volta results; the existing Volta community reports are 16 GB cards or mixed Volta+Pascal pairs, so a 32 GB pair at 512K may be a useful data point.

The headline is not a speed number — it is that decode tok/s on this engine is governed by tokens per verify window, which is a property of the content, not of the machine, and that the engine can now show you this directly with the profiler added in 0.1.40.

Hardware

  • 2× Tesla V100-PCIE-32GB (sm_70), passive, server-cooled. GPU0 ~29.5 GiB used, GPU1 ~32.2 GiB used (of 32,768 MiB each).
  • nvidia-smi topo -m: GPU0↔GPU1 = SYS — different CPU sockets, no P2P, no NVLink. Layer split (--layer-split) is the working path, as the docs describe.
  • CPU: 2× Xeon E5-2673 v3 (12 cores each, 48 threads total, AVX2 only — no AVX-512). PCIe Gen3 x16.
  • Memory: 125 GiB, DDR3-1600 (see the note at the end — this is a real shortcoming of this box and it turned out not to matter).
  • Linux, self-built engine with -DCMAKE_CUDA_ARCHITECTURES=70 -DSTRATA_EXPERIMENTAL_SM60=ON, CUDA 12.x.

Configuration

--max-context 524288 --rope-scaling yarn --rope-scale 2 --kv int8 --kv-resident 20480 --spec 4 --spec-min-p 0.70 --ple-io mmap --layer-split 20 --vision --vram-reserve-mib 700

Two notes on the reasoning, both measured rather than assumed:

  • --ple-io mmap, not direct or ram. ram is structurally wrong here, not just slower: locking the 26.8 GiB PLE table drives MemFree to 603 MB and the kernel then reclaims the process's own anonymous pages (VmSwap 1.45 GiB, pswpin +59%, compact_stall 324,567). The symptom is nasty because it does not look like a performance problem — the engine's own decode rate stays normal while a client sees tens of seconds of wall clock. direct pays an SSD read on every token's table lookup. mmap maps it into the page cache without locking, so rows in use stay resident and the rest stays reclaimable (VmLck 0, VmSwap 0). We are not quoting a percentage for mmap over direct — the arms were measured in separate windows and cannot be separated from run-to-run spread.
  • 512K is configured and works, with --kv-resident as a hard requirement rather than an optimisation. Without it the KV state takes the expert slots.

Decode: the profiler explains what moves the number

Setting STRATA_DECODE_TIMING=1 (with STRATA_VERIFY_PROFILE=1) logs a per-window breakdown per request. This is a genuinely useful addition — it is what let us stop guessing:

strata decode timing: 281 windows, avg T 2.04, 1.82 tokens/window, 25.80 ms/window
  = verify 23.56 (GPU-reach wait 0.00 + per-layer host 0.00 [plan 0.05 actq 0.07 jobs 0.00 CPU 0.51]
    + stage 0.06) + commit/emit 0.37 + draft 1.59
  per layer-window: CPU experts 0.10 (0.13 entries), VRAM hits 11.74, PCIe 0.01

Across seven profiled runs:

tokens/window ms/window ms/token CPU experts VRAM hits
3.09 37.89 12.26 0.50 18.99
2.79 33.49 12.00 0.41 17.08
2.55 33.19 13.02 0.35 16.46
1.84 25.80 14.02 0.20 12.18
1.59 24.06 15.13 0.22 9.86
1.50 23.58 15.72 0.12 9.41

ms/token is flat (12.0–15.7 ms) and ms/window is sub-linear in tokens/window — about 23 ms fixed per window plus ~9 ms for each token the verifier accepts. So throughput is set by tokens per window, which is how many draft tokens the verifier accepts on this content. The CPU expert pool is 0.5–4% of window time and PCIe is 0.01 ms: on this box neither is close to being the bottleneck, even at 512K with a DDR3-1600 CPU half.

Decode by content type — interleaved

Because a single number is not reproducible across content, here is the number with its content attached. 12 rounds per type, three types interleaved within one service lifetime (A,B,C,A,B,C…), first 2 rounds discarded, medians reported, and the whole thing run twice on separate occasions:

Content Round 1 (median) Round 2 (median) Acceptance
prose (zh) 84.8 80.8 78.8%
code 82.8 79.3 83.0%
prose (en) 70.9 68.3 81.5%
  • Content-to-content spread is 1.2×; the same content across the two rounds differs by ≤5%, and the ordering reproduced.
  • Acceptance rate does not predict tok/s — the lowest-acceptance arm (Chinese prose, 78.8%) was the fastest. What matters is tokens/window, i.e. acceptance combined with draft depth, not acceptance alone.
  • The ordering is not the intuitive one (we expected code to win). Worth measuring per workload rather than assuming.

Prefill (cold, engine counters)

Prompt Tokens read Engine prompt time Throughput
~256 315 1,736 ms 181.5 tok/s
~1K 1,004 1,892 ms 530.8 tok/s
~4K 3,707 3,425 ms 1,082.3 tok/s
~16K 14,572 8,653 ms 1,684.1 tok/s

Fresh prompt per row — a repeat is served from the conversation cache and reports a meaningless figure. Above 16K we could not measure cold without changing the configuration, so we do not claim those rows.

Other things that may be worth passing upstream

  • --spec 4 beat --spec 8 on this box (79.4 vs 72.9 tok/s at 256K, acceptance 76–78% vs 61–65%) — shallow drafts, higher hit rate. The --spec-min-p 0.70 row matched your calibration guidance.
  • --batch was a net loss here and the reason is in docs/BATCHING.md: batch windows carry no MTP drafts, and this configuration depends on MTP heavily. Four concurrent requests queued at 0.90× — consistent with "one request at a time".
  • --lookup-chain (opt-in, default 0): we measured this as a net loss on prose and code prompts (73.7 vs 77.0). Reading the help, it drafts by matching repeated context, so it should only pay on templated/repetitive content — which neither of our benchmark prompts is. Not a bug report, just a note that "no measurable gain" here is expected rather than surprising, and it might be worth a line in the help saying which workloads it targets.
  • --pool-affinity has no "restrict to one NUMA node" mode (only all/auto/p-cores). Wrapping the service in numactl --cpunodebind=0 was worse (49–69 tok/s), so we left the defaults alone. Noting it because "stop the workers crossing QPI" looks like an obvious win on a two-socket box and there is no supported way to try it.

What we did not measure

  • Full 512K recall. Our deepest retrieval test is 241,255 tokens (correct). 512K is configured and answers, but recall at the very top of the window is untested.
  • Quality benchmarks. No GSM8K/HumanEval/MMLU. Decode and prefill numbers here are safe to quote; quality numbers are absent.
  • A decode cost for 512K. We measured 256K and 512K and the difference sits inside the run-to-run spread, so we make no claim either way.
  • Long-run stability. These are figures from a live server, not a multi-day soak.

A note on the DDR3-1600 finding

This box runs DDR3-1600 although the E5-2673 v3 supports DDR4-2133, and docs/ warns that RAM below its rated speed slows the CPU half. It is a real shortcoming, and the profiler is what showed it does not matter here: the CPU expert pool is 0.5–4% of window time, so even doubling memory bandwidth has a few percent of headroom. Recording it because "check your RAM speed" is the intuitive first move on a box like this and the profiler says it is not where the time goes.

Full configuration, build notes and the complete negative-results log (including several of our own earlier claims that we had to withdraw) are in our write-up: https://github.com/ZackO2o/Strata-Qwen3.8-Flash-Next-512k-v100x2

Happy to re-run anything on request — the box is up and the profiler is one environment variable away.

Dominant language
C++
Stars
11.6k
Forks
1k
Avg merge
7h 46m
Merged PRs (30d)
30

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Niko1221/Strata

All issues in Niko1221/Strata

Similar issues

More C++ issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.