Community benchmark: IQ3_S and community AP-Q4_K_M — single RTX 5070 Ti vs 2× RTX 5060 Ti layer split
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 25/100
- Issue type
- Documentation
- Clarity
- Needs clarification
- Activity status
- Active
- Tech stack
- cpp
- Domain
- performance
Research direction
The issue provides benchmark configurations and results, but does not request a change or name a target file, test, or entry point. Ask maintainers what contribution they expect before starting; once scoped, define done around the requested artifact and any reproducibility checks.
Written by the indexing model from the issue text.
Description
Community benchmark: IQ3_S and community AP-Q4_K_M — single RTX 5070 Ti vs 2× RTX 5060 Ti layer split
Measured on 2026-10-03 by a ZCode-assisted lab operator. Tests Qwen3.8-Flash-Next IQ3_S
(GSQ-RCO) and the community agentionai AP-Q4_K_M GGUF on Strata 0.1.38, each on a single
RTX 5070 Ti (calibrated) and on a 2× RTX 5060 Ti layer split (calibrated) — four configurations
in total, same method throughout. Main limitations: 1–3 runs per point (decode of the native GSQ
path proved stable — ±5 % across repeats — while the imported GGML-format AP quant varies ±20 %+
with MTP luck), and 262,144-token prompts are refused by design (prompt + output exceeds the
262,144 context capacity).
Hardware and software
- 1× RTX 5070 Ti 16 GB (comparison config: 2× RTX 5060 Ti 16 GB, layer split); Intel Core Ultra 7 265K (20 cores, no AVX-512, AVX2 expert kernels); 96 GB (2×48 GB DDR5); PCIe 3.0 NVMe for models and the PLE table; single card on the CPU-attached x16 slot
- Ubuntu 24.04.4; driver 595.91.07; CUDA 13.0
- Strata commit
99f3dbd(2026-10-03); engine 0.1.38, source build - Background workloads: none during measurements (single-user lab machine)
Model and configuration
ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUFIQ3_S; community row:agentionai/Qwen3.8-Flash-Next-AP-GGUFAP-Q4_K_M, imported manually (iq_pack.py --compat-bf16 --experts-bin+ a BF16 one-tensor--embd-gguffor its Q6_K token embedding + its own mmproj-F16 wired through the config'svisionsection)- Vision encoder on for all four configurations
- Context 262,144; KV
int8(VRAM-resident;--kv-residentoff); expert cacheauto;--prefill auto - MTP
--spec 4 --spec-min-p 0.70; greedy sampling from the harness; calibrated per configuration: IQ3_S 5070 Ti--pcie-frac 0.20, IQ3_S split0.00, AP 5070 Ti0.35, AP split0.00(19 CPU workers each); speed projection off
# essence (IQ3_S, 5070 Ti; server 0.0.0.0:8080, gpu 0)
--pack packs/iq3_s --native Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf
--ple-gguf ...-00002-of-00002.gguf --expert-profile data/expert-profile.bin
--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.70 --pcie-frac 0.20
--mtp <data>/mtp/rt --max-context 262144 --kv int8 --vision
Method
Third-party OpenAI-compatible HTTP benchmark harness (GUI), one request at a time (the engine is
single-request serial), greedy (temperature 0), output cap 256 tokens per request, harness timeout
600 s for the longest prompts. Prompts are generated by the harness per length bucket (content
differs run to run within a bucket; re-running the same harness on a warm instance can inflate
decode via prompt-lookup drafting — noted below where it applies). Service freshly (re)started and
calibrated before each reported configuration; expert cache warm from one smoke request. Timing
boundaries are the harness's TTFT/ITL fields, cross-checked against the engine's per-request
decode_tok_s in /metrics (agrees within 2 %). Runs per point: IQ3_S split 2, IQ3_S 5070 Ti 1
(a second warm sweep landed within ±5 %), AP split 1 (after --prefill auto; the earlier
512-chunk sweep is excluded as a different configuration), AP 5070 Ti 3 (median reported; one of
the three sweeps ran on a warm instance and skews high). RAM observed: ~57 GB resident experts +
KV + runtime on IQ3_S; ~63 GB with AP-Q4_K_M.
Results
1. IQ3_S on 2× RTX 5060 Ti (layer split, calibrated)
| Actual prompt tokens | Reused tokens | Generated tokens | Runs | Prompt tok/s median and range | Decode tok/s median and range | TTFT seconds median and range |
|---|---|---|---|---|---|---|
| 512 | 0 | 256 | 2 | 542 (452–631) | 73.0 (64.4–81.7) | 0.91 (0.87–0.94) |
| 1024 | 0 | 256 | 2 | 798 (696–901) | 67.6 (61.7–73.5) | 1.11 (0.95–1.29) |
| 2048 | 0 | 256 | 2 | 1083 (963–1204) | 69.4 (54.5–84.3) | 1.48 (1.33–1.78) |
| 4096 | 0 | 256 | 2 | 1458 (1408–1508) | 70.4 (60.7–80.1) | 2.89 (2.80–2.99) |
| 8192 | 0 | 256 | 2 | 1602 (1560–1644) | 66.5 (64.0–69.1) | 5.21 (5.06–5.37) |
| 16384 | 0 | 256 | 2 | 2141 (2134–2149) | 65.2 (58.4–72.0) | 7.85 (7.72–7.98) |
| 32768 | 0 | 256 | 2 | 2552 (2520–2585) | 65.2 (60.2–70.1) | 13.29 (13.14–13.48) |
| 65536 | 0 | 256 | 2 | 2886 (2843–2929) | 68.5 (59.5–77.4) | 23.33 (23.23–23.42) |
| 131072 | 0 | 256 | 2 | 2969 (2905–3032) | 62.7 (54.3–71.1) | 45.17 (44.80–45.54) |
2. IQ3_S on RTX 5070 Ti (single card, calibrated)
| Actual prompt tokens | Reused tokens | Generated tokens | Runs | Prompt tok/s median and range | Decode tok/s median and range | TTFT seconds median and range |
|---|---|---|---|---|---|---|
| 512 | 0 | 256 | 1 | 845 | 118.3 | 0.66 |
| 1024 | 0 | 256 | 1 | 1325 | 100.2 | 0.83 |
| 2048 | 0 | 256 | 1 | 1982 | 100.4 | 1.10 |
| 4096 | 0 | 256 | 1 | 2822 | 100.3 | 1.52 |
| 8192 | 0 | 256 | 1 | 3241 | 100.4 | 2.60 |
| 16384 | 0 | 256 | 1 | 3373 | 99.4 | 4.97 |
| 32768 | 0 | 256 | 1 | 3393 | 95.9 | 9.80 |
| 65536 | 0 | 256 | 1 | 3333 | 99.5 | 19.9 |
| 131072 | 0 | 256 | 1 | 3147 | 99.1 | 42.0 |
IQ3_S, 5070 Ti vs split: decode +25–45 % at every length and essentially flat (95.9–100.4,
ITL ≈ 10 ms) vs the split's 63–73 with ~10 tok/s droop toward 131K; prefill ~2× in the 2–8K range,
+6 % at 131K; TTFT 45.2 → 42.0 s at 131K. The calibrated --pcie-frac flips 0.00 (split) → 0.20
(single).
3. AP-Q4_K_M on 2× RTX 5060 Ti (layer split, calibrated, --prefill auto)
| Actual prompt tokens | Reused tokens | Generated tokens | Runs | Prompt tok/s median and range | Decode tok/s median and range | TTFT seconds median and range |
|---|---|---|---|---|---|---|
| 512 | 0 | 256 | 1 | 452 | 64.4 | 1.21 |
| 1024 | 0 | 256 | 1 | 696 | 61.7 | 1.54 |
| 2048 | 0 | 256 | 1 | 963 | 54.5 | 2.19 |
| 4096 | 0 | 256 | 1 | 1408 | 60.7 | 2.99 |
| 8192 | 0 | 256 | 1 | 1560 | 64.0 | 5.34 |
| 16384 | 0 | 256 | 1 | 2134 | 58.4 | 7.98 |
| 32768 | 0 | 256 | 1 | 2585 | 60.2 | 13.44 |
| 65536 | 0 | 256 | 1 | 2929 | 59.5 | 23.42 |
| 131072 | 0 | 256 | 1 | 3032 | 54.3 | 43.57 |
4. AP-Q4_K_M on RTX 5070 Ti (single card, calibrated)
| Actual prompt tokens | Reused tokens | Generated tokens | Runs | Prompt tok/s median and range | Decode tok/s median and range | TTFT seconds median and range |
|---|---|---|---|---|---|---|
| 512 | 0 | 256 | 3 | 668 (636–679) | 92.1 (85.3–106.9) | 0.83 (0.81–0.86) |
| 1024 | 0 | 256 | 3 | 1129 (1034–1139) | 79.0 (76.1–83.3) | 0.97 (0.95–1.04) |
| 2048 | 0 | 256 | 3 | 1599 (1478–1606) | 77.1 (75.5–132.8) | 1.35 (1.33–1.44) |
| 4096 | 0 | 256 | 3 | 2758 (2448–2788) | 76.7 (75.2–76.9) | 1.56 (1.53–1.73) |
| 8192 | 0 | 256 | 3 | 2924 (2792–2917) | 114.4 (76.9–125.7) | 2.89 (2.89–3.01) |
| 16384 | 0 | 256 | 3 | 2987 (2931–3013) | 77.8 (74.9–82.6) | 5.58 (5.53–5.68) |
| 32768 | 0 | 256 | 3 | 2990 (2966–3016) | 84.0 (80.5–150.7) | 11.09 (10.99–11.17) |
| 65536 | 0 | 256 | 3 | 2922 (2920–2932) | 74.1 (72.1–96.7) | 22.61 (22.55–22.62) |
| 131072 | 0 | 256 | 3 | 2780 (2772–2795) | 76.9 (76.6–111.8) | 47.49 (47.31–47.63) |
AP-Q4_K_M, 5070 Ti vs split: decode +26–43 % at every length (medians 74–114 vs 54–64);
prefill +48–96 % below 16K, converging to ±0 % at 64K and −8 % at 131K (the only metric where the
single card does not win). --pcie-frac 0.00 → 0.35 (the imported GGML experts lean on the CPU
more, and the single fat card rewards a higher PCIe share).
Cross-model: on identical hardware the AP quant decodes ~20 % below IQ3_S (100.4 vs 79–77 at
mid lengths on the 5070 Ti) with ±20 %+ run-to-run variance vs the native path's ±5 %; per-request
MTP acceptance 50–63 % on fresh prompts vs 87 % for IQ3_S. Prefill is format-insensitive within
±5 % except at very long lengths.
Notes: 262,144-token prompts are refused by the server (prompt + 256 output exceeds the 262,144
capacity) — expected behavior, not an engine failure. The AP-Q4_K_M results are, to our
knowledge, the first published measurements of that community quant on Strata.
Correctness and limitations
- Sanity checks:
/health(context 262144, images on) before each sweep; one Chinese short-prompt
generation per configuration (correct, coherent); OCR of a rendered test image correct on IQ3_S
and on the imported AP quant (after wiring its own mmproj). - No needle/retrieval tests were run; quality is not assessed here, only speed.
- The AP-Q4_K_M import required
--compat-bf16, an--embd-ggufone-tensor BF16 file (its
token embedding is Q6_K, which the native path cannot dequantize on the GPU) and an
experts.binrepack for the full-RAM arena — described so others can reproduce. - Power limits: none applied; stock clocks. PCIe link idle-downgrades to Gen1 between requests
(observed), irrelevant to the measurements.
- Dominant language
- C++
- Stars
- 11.6k
- Forks
- 1k
- Avg merge
- 7h 46m
- Merged PRs (30d)
- 30
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Niko1221/Strata
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Niko1221/Strata#974 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
Similar issues
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Qiskit/qiskit-aer#2466 ·
-
feature request
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 2 days
-
status:needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
PX4/PX4-Autopilot#29006 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 1 day
-
Segfault in pwStreamAddBuffer: createBuffer() returning nullptr is dereferenced (Screencopy.cpp:943)Open
Difficulty 2/5 1-3 hours Newbie friendliness 74/100