Community result: 2.5-3.1x output at 256K context on a 16 GB card (IQ3_S) - calibration, --spec 6, learned profile
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Newbie friendliness
- 55/100
Research direction
Start by reading the --setup handling in START-HERE.bat and how it writes the config; compare that with the settings-file behavior described in the issue. Decide whether re-running setup should preserve the --expert-profile/--expert-profile-save pair or print a note, then verify the chosen behavior.
Written by the indexing model from the issue text.
Description
Sharing a measured result and a writeup: docs/TUNING-256K-16GB.md on my fork ->
https://github.com/JiuYue0820/Strata/blob/docs-256k-tuning/docs/TUNING-256K-16GB.md
PC: RTX 5070 Ti 16 GB, i7-14700KF (8P+12E), 96 GB DDR5-5600, Windows, engine 0.1.38, Qwen3.8-Flash-Next IQ3_S (GSQ-RCO), vision encoder on GPU, context 262,144, KV int8 + streaming (--kv-resident 32768).
Changes: START-HERE.bat --calibrate (kept --pool-workers 13, --spec-min-p 0.70, --pcie-frac 0.00), --spec 6, and the learned expert profile (--expert-profile-save).
Benchmark: one 257,630-token prompt (this repo's own source/docs) through the HTTP API, greedy, reasoning_effort: none. "Warm" = prefix reused, which is how agent sessions actually run. Single runs, so the usual ±20% noise applies.
| Configuration | Cold | Warm |
|---|---|---|
Stock (--spec 4, no calibration, shipped profile) |
17.5 tok/s | 17.2 tok/s |
--spec 6 + calibrated |
43.0 tok/s | 53.5 tok/s |
| + learned expert profile | 45.5 tok/s | 47.0 tok/s |
Prompt reading went from 1,468 to 1,744 tok/s.
The calibration internals on this machine, in case they are useful data points:
--pool-workers: 19 → 36.9 tok/s, 13 → 54.4, 10 → 41.3 (E-cores stall every verify window; the defaults were measured on a 6-core Ryzen with no E-cores)--spec-min-p: 0.3 → 36.6, 0.5 → 41.7, 0.7 → 44.6--pcie-frac: 0.0 → 42.1, 0.2 → 37.9, 0.35 → 27.9, 0.55 → 23.9, 0.75 → 20.7 (PCIe 5 x16 + fast CPU: computing a missed expert beats copying it)
Also fills in the IQ3_S 262K row that DETAILS.md leaves unmeasured: on 96 GB of RAM it runs with room to spare, 38.8-53.5 tok/s across all post-tuning runs.
One documentation gap this exposed: a --setup re-run rewrites the config and silently drops a manually added --expert-profile/--expert-profile-save pair (the three calibrated settings survive via the settings file). Worth either preserving or printing a note.
- Dominant language
- C++
- Stars
- 11.6k
- Forks
- 1k
- Avg merge
- 7h 46m
- Merged PRs (30d)
- 30
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Niko1221/Strata
-
expert_cache_segmented_test fails on HIP builds instead of skipping (--vram-elastic is CUDA-only)Possibly taken A pull request linked to this issue is open or already merged. Open
Difficulty 2/5 1-3 hours Newbie friendliness 83/100
Maintainers usually reply within 1 day
-
hip_q2_zero fails on gfx1201 (R9700) with ROCm 7.10: Q2_0 signed-zero fix e9a5f8d is gated to gfx1012 / HIP < 7Possibly taken A pull request linked to this issue is open or already merged. Open
Difficulty 2/5 1-3 hours Newbie friendliness 66/100
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 82/100
Maintainers usually reply within 1 day
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
? - Needs Triage bot_watch bug
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
NVIDIA/cudf-spark-jni#5267 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
(S1 - Need confirmation)
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
CleverRaven/Cataclysm-DDA#88974 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
microsoft/onnxruntime#33215 ·
Maintainers usually reply within 2 days