Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Community benchmark: IQ3_S and community AP-Q4_K_M — single RTX 5070 Ti vs 2× RTX 5060 Ti layer split

Open
#791 1 comment 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
25/100
Issue type
Documentation
Clarity
Needs clarification
Activity status
Active
Tech stack
cpp
Domain
performance

Research direction

The issue provides benchmark configurations and results, but does not request a change or name a target file, test, or entry point. Ask maintainers what contribution they expect before starting; once scoped, define done around the requested artifact and any reproducibility checks.

Written by the indexing model from the issue text.

Description

Community benchmark: IQ3_S and community AP-Q4_K_M — single RTX 5070 Ti vs 2× RTX 5060 Ti layer split

Measured on 2026-10-03 by a ZCode-assisted lab operator. Tests Qwen3.8-Flash-Next IQ3_S
(GSQ-RCO) and the community agentionai AP-Q4_K_M GGUF on Strata 0.1.38, each on a single
RTX 5070 Ti (calibrated) and on a 2× RTX 5060 Ti layer split (calibrated) — four configurations
in total, same method throughout. Main limitations: 1–3 runs per point (decode of the native GSQ
path proved stable — ±5 % across repeats — while the imported GGML-format AP quant varies ±20 %+
with MTP luck), and 262,144-token prompts are refused by design (prompt + output exceeds the
262,144 context capacity).

Hardware and software

  • 1× RTX 5070 Ti 16 GB (comparison config: 2× RTX 5060 Ti 16 GB, layer split); Intel Core Ultra 7 265K (20 cores, no AVX-512, AVX2 expert kernels); 96 GB (2×48 GB DDR5); PCIe 3.0 NVMe for models and the PLE table; single card on the CPU-attached x16 slot
  • Ubuntu 24.04.4; driver 595.91.07; CUDA 13.0
  • Strata commit 99f3dbd (2026-10-03); engine 0.1.38, source build
  • Background workloads: none during measurements (single-user lab machine)

Model and configuration

  • ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF IQ3_S; community row: agentionai/Qwen3.8-Flash-Next-AP-GGUF AP-Q4_K_M, imported manually (iq_pack.py --compat-bf16 --experts-bin + a BF16 one-tensor --embd-gguf for its Q6_K token embedding + its own mmproj-F16 wired through the config's vision section)
  • Vision encoder on for all four configurations
  • Context 262,144; KV int8 (VRAM-resident; --kv-resident off); expert cache auto; --prefill auto
  • MTP --spec 4 --spec-min-p 0.70; greedy sampling from the harness; calibrated per configuration: IQ3_S 5070 Ti --pcie-frac 0.20, IQ3_S split 0.00, AP 5070 Ti 0.35, AP split 0.00 (19 CPU workers each); speed projection off
# essence (IQ3_S, 5070 Ti; server 0.0.0.0:8080, gpu 0)
--pack packs/iq3_s --native Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf
--ple-gguf ...-00002-of-00002.gguf --expert-profile data/expert-profile.bin
--expert-cache auto --prefill auto --spec 4 --spec-min-p 0.70 --pcie-frac 0.20
--mtp <data>/mtp/rt --max-context 262144 --kv int8 --vision

Method

Third-party OpenAI-compatible HTTP benchmark harness (GUI), one request at a time (the engine is
single-request serial), greedy (temperature 0), output cap 256 tokens per request, harness timeout
600 s for the longest prompts. Prompts are generated by the harness per length bucket (content
differs run to run within a bucket; re-running the same harness on a warm instance can inflate
decode via prompt-lookup drafting — noted below where it applies). Service freshly (re)started and
calibrated before each reported configuration; expert cache warm from one smoke request. Timing
boundaries are the harness's TTFT/ITL fields, cross-checked against the engine's per-request
decode_tok_s in /metrics (agrees within 2 %). Runs per point: IQ3_S split 2, IQ3_S 5070 Ti 1
(a second warm sweep landed within ±5 %), AP split 1 (after --prefill auto; the earlier
512-chunk sweep is excluded as a different configuration), AP 5070 Ti 3 (median reported; one of
the three sweeps ran on a warm instance and skews high). RAM observed: ~57 GB resident experts +
KV + runtime on IQ3_S; ~63 GB with AP-Q4_K_M.

Results

1. IQ3_S on 2× RTX 5060 Ti (layer split, calibrated)
Actual prompt tokens Reused tokens Generated tokens Runs Prompt tok/s median and range Decode tok/s median and range TTFT seconds median and range
512 0 256 2 542 (452–631) 73.0 (64.4–81.7) 0.91 (0.87–0.94)
1024 0 256 2 798 (696–901) 67.6 (61.7–73.5) 1.11 (0.95–1.29)
2048 0 256 2 1083 (963–1204) 69.4 (54.5–84.3) 1.48 (1.33–1.78)
4096 0 256 2 1458 (1408–1508) 70.4 (60.7–80.1) 2.89 (2.80–2.99)
8192 0 256 2 1602 (1560–1644) 66.5 (64.0–69.1) 5.21 (5.06–5.37)
16384 0 256 2 2141 (2134–2149) 65.2 (58.4–72.0) 7.85 (7.72–7.98)
32768 0 256 2 2552 (2520–2585) 65.2 (60.2–70.1) 13.29 (13.14–13.48)
65536 0 256 2 2886 (2843–2929) 68.5 (59.5–77.4) 23.33 (23.23–23.42)
131072 0 256 2 2969 (2905–3032) 62.7 (54.3–71.1) 45.17 (44.80–45.54)
2. IQ3_S on RTX 5070 Ti (single card, calibrated)
Actual prompt tokens Reused tokens Generated tokens Runs Prompt tok/s median and range Decode tok/s median and range TTFT seconds median and range
512 0 256 1 845 118.3 0.66
1024 0 256 1 1325 100.2 0.83
2048 0 256 1 1982 100.4 1.10
4096 0 256 1 2822 100.3 1.52
8192 0 256 1 3241 100.4 2.60
16384 0 256 1 3373 99.4 4.97
32768 0 256 1 3393 95.9 9.80
65536 0 256 1 3333 99.5 19.9
131072 0 256 1 3147 99.1 42.0

IQ3_S, 5070 Ti vs split: decode +25–45 % at every length and essentially flat (95.9–100.4,
ITL ≈ 10 ms) vs the split's 63–73 with ~10 tok/s droop toward 131K; prefill ~2× in the 2–8K range,
+6 % at 131K; TTFT 45.2 → 42.0 s at 131K. The calibrated --pcie-frac flips 0.00 (split) → 0.20
(single).

3. AP-Q4_K_M on 2× RTX 5060 Ti (layer split, calibrated, --prefill auto)
Actual prompt tokens Reused tokens Generated tokens Runs Prompt tok/s median and range Decode tok/s median and range TTFT seconds median and range
512 0 256 1 452 64.4 1.21
1024 0 256 1 696 61.7 1.54
2048 0 256 1 963 54.5 2.19
4096 0 256 1 1408 60.7 2.99
8192 0 256 1 1560 64.0 5.34
16384 0 256 1 2134 58.4 7.98
32768 0 256 1 2585 60.2 13.44
65536 0 256 1 2929 59.5 23.42
131072 0 256 1 3032 54.3 43.57
4. AP-Q4_K_M on RTX 5070 Ti (single card, calibrated)
Actual prompt tokens Reused tokens Generated tokens Runs Prompt tok/s median and range Decode tok/s median and range TTFT seconds median and range
512 0 256 3 668 (636–679) 92.1 (85.3–106.9) 0.83 (0.81–0.86)
1024 0 256 3 1129 (1034–1139) 79.0 (76.1–83.3) 0.97 (0.95–1.04)
2048 0 256 3 1599 (1478–1606) 77.1 (75.5–132.8) 1.35 (1.33–1.44)
4096 0 256 3 2758 (2448–2788) 76.7 (75.2–76.9) 1.56 (1.53–1.73)
8192 0 256 3 2924 (2792–2917) 114.4 (76.9–125.7) 2.89 (2.89–3.01)
16384 0 256 3 2987 (2931–3013) 77.8 (74.9–82.6) 5.58 (5.53–5.68)
32768 0 256 3 2990 (2966–3016) 84.0 (80.5–150.7) 11.09 (10.99–11.17)
65536 0 256 3 2922 (2920–2932) 74.1 (72.1–96.7) 22.61 (22.55–22.62)
131072 0 256 3 2780 (2772–2795) 76.9 (76.6–111.8) 47.49 (47.31–47.63)

AP-Q4_K_M, 5070 Ti vs split: decode +26–43 % at every length (medians 74–114 vs 54–64);
prefill +48–96 % below 16K, converging to ±0 % at 64K and −8 % at 131K (the only metric where the
single card does not win). --pcie-frac 0.00 → 0.35 (the imported GGML experts lean on the CPU
more, and the single fat card rewards a higher PCIe share).

Cross-model: on identical hardware the AP quant decodes ~20 % below IQ3_S (100.4 vs 79–77 at
mid lengths on the 5070 Ti) with ±20 %+ run-to-run variance vs the native path's ±5 %; per-request
MTP acceptance 50–63 % on fresh prompts vs 87 % for IQ3_S. Prefill is format-insensitive within
±5 % except at very long lengths.

Notes: 262,144-token prompts are refused by the server (prompt + 256 output exceeds the 262,144
capacity) — expected behavior, not an engine failure. The AP-Q4_K_M results are, to our
knowledge, the first published measurements of that community quant on Strata.

Correctness and limitations

  • Sanity checks: /health (context 262144, images on) before each sweep; one Chinese short-prompt
    generation per configuration (correct, coherent); OCR of a rendered test image correct on IQ3_S
    and on the imported AP quant (after wiring its own mmproj).
  • No needle/retrieval tests were run; quality is not assessed here, only speed.
  • The AP-Q4_K_M import required --compat-bf16, an --embd-gguf one-tensor BF16 file (its
    token embedding is Q6_K, which the native path cannot dequantize on the GPU) and an
    experts.bin repack for the full-RAM arena — described so others can reproduce.
  • Power limits: none applied; stock clocks. PCIe link idle-downgrades to Gen1 between requests
    (observed), irrelevant to the measurements.
Dominant language
C++
Stars
11.6k
Forks
1k
Avg merge
7h 46m
Merged PRs (30d)
30

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Niko1221/Strata

All issues in Niko1221/Strata

Similar issues

More C++ issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.