Local-adaptation tuning directions (method, not values) — plus the hardware-specialization layer (model → GPU → CPU) that I'm building
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 15/100
- Tipo di issue
- Documentazione
- Chiarezza
- Da chiarire
- Stato di attività
- Attiva
- Stack tecnologico
- cpp
- Ambito
- documentation
Direzione di ricerca
The issue mentions DETAILS.md, --calibrate, and bench/results/, but asks whether and where to contribute rather than defining an accepted change. First read the existing tuning guidance and check with the maintainers about scope and placement; the issue says findings should come before document or code changes. Done criteria are not established until the maintainers decide what they want.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Hi — I'm @1314521gjy. I adapted this engine to my machine (RTX 4080S 32 GB + 96 GB), and that work's report is already merged (#780); #1351 (a report) and #1354 (docs) are still open. Before this I did the same kind of work on another local stack — a hardware-adaptation fusion build at github.com/1314521gjy/ninfer-fusion-kvmem — so what I do is make an engine fit a specific card and CPU.
This issue brings two things: a set of tuning directions (method, not values, from roughly a hundred controlled runs on one machine), and the hardware-specialization layer I am working on now — this issue is one of that work's outputs, not a suggestion I'm handing over.
What I'd like to add
Your --calibrate and DETAILS.md already say which knobs a PC should tune. I'd like to contribute the other half: how to tune without lying to yourself. It deliberately carries directions, not values — on my own machine one unchanged config measured 111 → 134 tok/s across sessions, so values are not portable.
The rules that changed our results:
- Self-witness first. Every knob needs a field the engine prints itself, that can go red. Without it, "no difference" is unreadable — the switch may simply not have engaged. We wasted a full sweep of arms on a switch nobody could prove had engaged, which is what made this a rule.
- Discard a warm-up arm, and mirror the order (A B B A). The first arm of a session is systematically cold, and absolute speed drifts; only same-batch comparisons are evidence.
- The VRAM-headroom cliff. When free VRAM approaches zero the driver pages, the engine reports nothing, and timing stops being predictable rather than merely slower: we measured the same deterministic counters 45% apart. The field to watch is the one the engine prints when it captures each window — not the cache size you configured.
- Some knobs are machine-specific, and we have a second machine to prove it. A member of our group ran the same directions on a 24 GB laptop card with a hybrid-core CPU and found the opposite optima for the CPU expert thread count and for the CPU/GPU miss split. His write-up (shared with his permission): context 256K → 512K at almost no cost, expert slots +26%, cold prefill +13–19%, and he dodged the headroom cliff — 22.8 → 66.0 tok/s once free VRAM went from 262 MiB to 1069 MiB. He labels that explicitly as dodging a trap, not tuning. Same method, different machine, opposite values: the method transfers, the values do not.
- A feature that is a big win elsewhere can be worth nothing here. The elastic-KV route that gave him +29% expert slots gave us no slot gain at all — because we already run KV streaming, which solves the same problem. Same mechanism, opposite conclusion.
- Don't use the acceptance rate as a verdict. Tightening the draft floor raised acceptance (0.60 → 0.80) and still got slower; what matters is tokens per verification, and then wall-clock.
What I'm working on: completing the specialization at the hardware layer (model → GPU → CPU)
This part is current work on my side, and this issue is one of its outputs — not a suggestion I'm handing over.
This engine is a model-specific weapon: one model's geometry is hard-wired, and that hard-wiring is where its order-of-magnitude advantage comes from. What I'm working on is completing the same idea in the other direction — giving the hardware the same treatment:
- GPU-specific: how much VRAM is left on this card, how wide is this link, and what its bandwidth and latency actually are.
- CPU-specific: how many cores, which instruction set, how many memory channels, and how the host thread and the worker pool are laid out.
- Both treated as first-class inputs rather than one-off calibration — and both verifiable: every profile binds to a field the engine prints.
The model layer answers for whom; the hardware layers answer on what. The light route I'm using: probe → pick a profile → bind each profile to a self-witness field → keep profile + readings + rollback as one portable record. Then the third layer is not guesswork — it is checkable the way the first layer is.
The direction guide above is the first product of that work; the profile-and-record side is next. I'd rather show findings first, and touch your tree only when you say so.
Three questions — any answer is fine, including "no"
- Would you like this as a document in the repo? And if so, where —
docs/, a section ofDETAILS.md, orbench/results/? If you'd rather not, say so and I'll keep it out of the repo. - Is the hardware-specialization layer useful to you — and would you like the outputs of that work to keep coming back to you in this shape (findings first, doc/code changes only on your word)?
- May I add one line to the tuning part of
DETAILS.md— "always measure, and bind each knob to a field the engine prints" — as a minimal separate PR?
No AI-assisted wording anywhere; I sign as 1314521gjy.
- Lingua principale
- C++
- Stelle
- 11.6k
- Fork
- 1k
- Merge medio
- 7h 46m
- PR unite (30g)
- 30
Preparare l'ambiente
Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di Niko1221/Strata
-
expert_cache_segmented_test fails on HIP builds instead of skipping (--vram-elastic is CUDA-only)Forse già presa Una pull request collegata a questa issue è aperta o già unita. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 83/100
I maintainer di solito rispondono entro 1 giorno
-
hip_q2_zero fails on gfx1201 (R9700) with ROCm 7.10: Q2_0 signed-zero fix e9a5f8d is gated to gfx1012 / HIP < 7Forse già presa Una pull request collegata a questa issue è aperta o già unita. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 66/100
I maintainer di solito rispondono entro 1 giorno
-
MI50 32 GB (gfx906) on 0.1.40.1: 126K/252K needles, a 16 GB-limit run, temperatures (results)Aperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 82/100
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di Niko1221/Strata
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno
-
? - Needs Triage bot_watch bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
NVIDIA/cudf-spark-jni#5267 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
I maintainer di solito rispondono entro 1 giorno
-
(S1 - Need confirmation)
Difficoltà 2/5 1-3 ore Idoneità per principianti 62/100
CleverRaven/Cataclysm-DDA#88974 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 62/100
microsoft/onnxruntime#33215 ·
I maintainer di solito rispondono entro 2 giorni