Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Community result: 2.5-3.1x output at 256K context on a 16 GB card (IQ3_S) - calibration, --spec 6, learned profile

Closed
#775 1 comment 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
3/5
Estimated time
1-2 days
Newbie friendliness
55/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
cpp, shell
Domain
tooling

Research direction

Start by reading the --setup handling in START-HERE.bat and how it writes the config; compare that with the settings-file behavior described in the issue. Decide whether re-running setup should preserve the --expert-profile/--expert-profile-save pair or print a note, then verify the chosen behavior.

Written by the indexing model from the issue text.

Description

Sharing a measured result and a writeup: docs/TUNING-256K-16GB.md on my fork ->
https://github.com/JiuYue0820/Strata/blob/docs-256k-tuning/docs/TUNING-256K-16GB.md

PC: RTX 5070 Ti 16 GB, i7-14700KF (8P+12E), 96 GB DDR5-5600, Windows, engine 0.1.38, Qwen3.8-Flash-Next IQ3_S (GSQ-RCO), vision encoder on GPU, context 262,144, KV int8 + streaming (--kv-resident 32768).

Changes: START-HERE.bat --calibrate (kept --pool-workers 13, --spec-min-p 0.70, --pcie-frac 0.00), --spec 6, and the learned expert profile (--expert-profile-save).

Benchmark: one 257,630-token prompt (this repo's own source/docs) through the HTTP API, greedy, reasoning_effort: none. "Warm" = prefix reused, which is how agent sessions actually run. Single runs, so the usual ±20% noise applies.

Configuration Cold Warm
Stock (--spec 4, no calibration, shipped profile) 17.5 tok/s 17.2 tok/s
--spec 6 + calibrated 43.0 tok/s 53.5 tok/s
+ learned expert profile 45.5 tok/s 47.0 tok/s

Prompt reading went from 1,468 to 1,744 tok/s.

The calibration internals on this machine, in case they are useful data points:

  • --pool-workers: 19 → 36.9 tok/s, 13 → 54.4, 10 → 41.3 (E-cores stall every verify window; the defaults were measured on a 6-core Ryzen with no E-cores)
  • --spec-min-p: 0.3 → 36.6, 0.5 → 41.7, 0.7 → 44.6
  • --pcie-frac: 0.0 → 42.1, 0.2 → 37.9, 0.35 → 27.9, 0.55 → 23.9, 0.75 → 20.7 (PCIe 5 x16 + fast CPU: computing a missed expert beats copying it)

Also fills in the IQ3_S 262K row that DETAILS.md leaves unmeasured: on 96 GB of RAM it runs with room to spare, 38.8-53.5 tok/s across all post-tuning runs.

One documentation gap this exposed: a --setup re-run rewrites the config and silently drops a manually added --expert-profile/--expert-profile-save pair (the three calibrated settings survive via the settings file). Worth either preserving or printing a note.

Dominant language
C++
Stars
11.6k
Forks
1k
Avg merge
7h 46m
Merged PRs (30d)
30

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Niko1221/Strata

All issues in Niko1221/Strata

Similar issues

More C++ issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.