docs: `--lookup-chain` — note the workload it targets (repeating context), since the default suffix path already covers non-repeating text
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 1/5
- Estimated time
- Under an hour
- Newbie friendliness
- 82/100
- Issue type
- Documentation
- Clarity
- Clearly specified
- Activity status
- Active
- Domain
- documentation
Research direction
Locate the help/usage entry that prints the --lookup-chain K line (the issue quotes it verbatim) and the sibling --suffix-draft line for tone; generate.cpp:2089 is the mechanism reference showing the default suffix path. Add the suggested phrase noting the setting pays when the context repeats (templated output, fixed-format turns, repeated code blocks), then rebuild or run the binary's help output to confirm the line renders within its column alignment. Done when the help text reads like the --suffix-draft wording and no other code changes.
Written by the indexing model from the issue text.
Description
What this is: a short measurement note on --lookup-chain, plus a documentation observation. Not a bug report.
Hardware: 2× Tesla V100-PCIE-32GB (sm_70, CUDA 12 source build), Qwen3.8-Flash-Next IQ3_S, 512K, --layer-split 20, --spec 4 --spec-min-p 0.70. Same box as the community report in #1300.
What we measured
We swept three opt-in settings, 12 runs each, first 2 discarded, medians, arms interleaved within one service lifetime:
| Setting | prose | code | vs default |
|---|---|---|---|
| default | 77.0 | 103.1 | — |
--lookup-chain 2 |
76.1 | 99.8 | -1.2% / -3.2% |
--lookup-chain 4 |
73.7 | 100.8 | -4.3% / -2.2% |
--host-core last |
76.3 | 101.2 | -0.9% / -1.8% |
--host-core last we understand — its help says the win comes from Windows sending a GPU's interrupts to one logical processor, which is a Windows-only condition; on Linux there is nothing to fix. Correctly documented, no action needed.
--lookup-chain is the one worth a sentence in the help. Reading the source (generate.cpp:2089), the default path already enables prompt lookup:
// Prompt lookup (the suffix drafter, on by default): the MTP keeps its --spec windows and a lookup window may be
// up to 2 tokens longer; the draft policy (strata/spec/draft_policy.hpp) takes one only where it pays. Code
// edits +6-11%, ordinary text unchanged (bench/results/2026-09-27-spec). --suffix-draft 0 turns it off.
if (o.suffix_draft > 0 && o.spec >= 2 && o.mtp_max_t == 0) {
o.mtp_max_t = o.spec;
o.spec = std::min(o.spec + 2, 8);
}
So the default already takes the suffix path when the draft policy says it pays, and the +6–11% case documented there is code edits. --lookup-chain K is specifically "add up to K prompt-lookup drafts that continue the MTP's drafts", which needs the context to actually repeat itself — our prose and code prompts do not, so a null-to-negative result is the expected outcome rather than a surprising one.
The suggestion
The help entry describes the mechanism precisely but not the workload it targets:
--lookup-chain K opt-in: after the MTP's drafts, add up to K prompt-lookup drafts that continue
them (the window grows to at most 8; default 0 = off)
It is opt-in and off by default, so nobody loses anything by it — but "opt-in" invites trying it, and the setting that decides the outcome (does your context repeat?) is not mentioned but is mentioned on the sibling --suffix-draft line ("when it pays"). A phrase in the same spirit on the --lookup-chain line — "pays when the context repeats (templated output, fixed-format turns, code with repeated blocks)" — would save the next person the sweep we just did.
What we did not establish
- We have not tested
--lookup-chainon a templated workload, so this is not a claim that it never helps — only that on non-repeating prose/code it does not, and that that is consistent with the mechanism. - Our sweep ran in one batch each; the numbers above are the within-batch comparison, and we would not quote the absolute values across batches (this box has a documented content-dependence in decode rate — see #1300).
- Dominant language
- C++
- Stars
- 11.6k
- Forks
- 1k
- Avg merge
- 7h 46m
- Merged PRs (30d)
- 30
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Niko1221/Strata
-
expert_cache_segmented_test fails on HIP builds instead of skipping (--vram-elastic is CUDA-only)Open
Difficulty 2/5 1-3 hours Newbie friendliness 83/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 66/100
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
Maintainers usually reply within 1 day
Similar issues
-
Difficulty 1/5 Under an hour Newbie friendliness 78/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
EsotericSoftware/spine-runtimes#3186 ·
-
An empty line splits a signature where an ordinary comment is right above an argument's HaddockOpen
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
mrkkrp/tilia#213 · 1 comment ·
Maintainers usually reply within 1 day
-
Round video messages start gray and blocky with libx264: encoder is configured for 1,000,000 fpsOpen
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
telegramdesktop/tdesktop#31422 ·
Maintainers usually reply within 9 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
Maintainers usually reply within 5 days