Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Local-adaptation tuning directions (method, not values) — plus the hardware-specialization layer (model → GPU → CPU) that I'm building

Aperta
#1,483 2 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
5/5
Tempo stimato
Più di una settimana
Idoneità per principianti
15/100
Tipo di issue
Documentazione
Chiarezza
Da chiarire
Stato di attività
Attiva
Stack tecnologico
cpp
Ambito
documentation

Direzione di ricerca

The issue mentions DETAILS.md, --calibrate, and bench/results/, but asks whether and where to contribute rather than defining an accepted change. First read the existing tuning guidance and check with the maintainers about scope and placement; the issue says findings should come before document or code changes. Done criteria are not established until the maintainers decide what they want.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Hi — I'm @1314521gjy. I adapted this engine to my machine (RTX 4080S 32 GB + 96 GB), and that work's report is already merged (#780); #1351 (a report) and #1354 (docs) are still open. Before this I did the same kind of work on another local stack — a hardware-adaptation fusion build at github.com/1314521gjy/ninfer-fusion-kvmem — so what I do is make an engine fit a specific card and CPU.

This issue brings two things: a set of tuning directions (method, not values, from roughly a hundred controlled runs on one machine), and the hardware-specialization layer I am working on now — this issue is one of that work's outputs, not a suggestion I'm handing over.

What I'd like to add

Your --calibrate and DETAILS.md already say which knobs a PC should tune. I'd like to contribute the other half: how to tune without lying to yourself. It deliberately carries directions, not values — on my own machine one unchanged config measured 111 → 134 tok/s across sessions, so values are not portable.

The rules that changed our results:

  • Self-witness first. Every knob needs a field the engine prints itself, that can go red. Without it, "no difference" is unreadable — the switch may simply not have engaged. We wasted a full sweep of arms on a switch nobody could prove had engaged, which is what made this a rule.
  • Discard a warm-up arm, and mirror the order (A B B A). The first arm of a session is systematically cold, and absolute speed drifts; only same-batch comparisons are evidence.
  • The VRAM-headroom cliff. When free VRAM approaches zero the driver pages, the engine reports nothing, and timing stops being predictable rather than merely slower: we measured the same deterministic counters 45% apart. The field to watch is the one the engine prints when it captures each window — not the cache size you configured.
  • Some knobs are machine-specific, and we have a second machine to prove it. A member of our group ran the same directions on a 24 GB laptop card with a hybrid-core CPU and found the opposite optima for the CPU expert thread count and for the CPU/GPU miss split. His write-up (shared with his permission): context 256K → 512K at almost no cost, expert slots +26%, cold prefill +13–19%, and he dodged the headroom cliff — 22.8 → 66.0 tok/s once free VRAM went from 262 MiB to 1069 MiB. He labels that explicitly as dodging a trap, not tuning. Same method, different machine, opposite values: the method transfers, the values do not.
  • A feature that is a big win elsewhere can be worth nothing here. The elastic-KV route that gave him +29% expert slots gave us no slot gain at all — because we already run KV streaming, which solves the same problem. Same mechanism, opposite conclusion.
  • Don't use the acceptance rate as a verdict. Tightening the draft floor raised acceptance (0.60 → 0.80) and still got slower; what matters is tokens per verification, and then wall-clock.

What I'm working on: completing the specialization at the hardware layer (model → GPU → CPU)

This part is current work on my side, and this issue is one of its outputs — not a suggestion I'm handing over.

This engine is a model-specific weapon: one model's geometry is hard-wired, and that hard-wiring is where its order-of-magnitude advantage comes from. What I'm working on is completing the same idea in the other direction — giving the hardware the same treatment:

  • GPU-specific: how much VRAM is left on this card, how wide is this link, and what its bandwidth and latency actually are.
  • CPU-specific: how many cores, which instruction set, how many memory channels, and how the host thread and the worker pool are laid out.
  • Both treated as first-class inputs rather than one-off calibration — and both verifiable: every profile binds to a field the engine prints.

The model layer answers for whom; the hardware layers answer on what. The light route I'm using: probe → pick a profile → bind each profile to a self-witness field → keep profile + readings + rollback as one portable record. Then the third layer is not guesswork — it is checkable the way the first layer is.

The direction guide above is the first product of that work; the profile-and-record side is next. I'd rather show findings first, and touch your tree only when you say so.

Three questions — any answer is fine, including "no"

  1. Would you like this as a document in the repo? And if so, where — docs/, a section of DETAILS.md, or bench/results/? If you'd rather not, say so and I'll keep it out of the repo.
  2. Is the hardware-specialization layer useful to you — and would you like the outputs of that work to keep coming back to you in this shape (findings first, doc/code changes only on your word)?
  3. May I add one line to the tuning part of DETAILS.md — "always measure, and bind each knob to a field the engine prints" — as a minimal separate PR?

No AI-assisted wording anywhere; I sign as 1314521gjy.

Lingua principale
C++
Stelle
11.6k
Fork
1k
Merge medio
7h 46m
PR unite (30g)
30

Preparare l'ambiente

Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di Niko1221/Strata

Tutte le issue di Niko1221/Strata

Issue simili

Altre issue su C++

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.