Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

P40 + 3070 (mixed Pascal/Ampere pair): `layer_split: auto` vs forced splits on 0.1.39, with hit rates (follow-up to #604)

Aperta
#875 2 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
35/100
Tipo di issue
Funzionalità
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
cpp

Direzione di ricerca

Start with the related placement-search and weight-carve issues (#584, #639, and #559), then inspect how the layer-split cost model and auto selection account for trimming. The issue names bench/results/.../strata-bench.py and provides configs and engine logs for comparison. Done means determining whether auto can use trimming and validating the resulting split and performance against the reported forced-split runs.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

p40-3070-logs-v0.1.39.tar.gz

Follow-up to #604, where you asked for "a P40 + 3070 with layer_split: auto against a few forced splits (--stats and the hit rates)". Here is that run, on v0.1.39 as released (no patches; CUDA 12.9 engine, STRATA_EXPERIMENTAL_SM60=1), IQ3_XXS, 36,864-token context, P40 power-capped at 130 W, DDR4-2400, PCIe 3.0 x8. Three runs per cell with bench/results/.../strata-bench.py (unique prefix per run, 256 tokens). Prefill / decode in tok/s; hit = decode expert-cache hit rate range over the runs.

config (3070 first unless noted) slots 3070 / P40 chunk 4K 32K hit
P40 alone - / 10,451 8192 322 / 34 339 / 34 83-96%
3070 alone 513 / - 1024 144 / 30 155 / 37 14-29%
auto (K=43) 1,536 / 2,560 4096 462 / 35 617 / 42 39-58%
auto, P40 first 8,192 / 736 1024 193 / 32 221 / 37 26-42%
forced K=8 2,177 / 10,333 5888 333 / 39 362 / 37 89-98%
forced K=20 2,070 / 9,780 5120 386 / 38 491 / 38 82-89%
K=20 + STRATA_STAGE_TRIM=1 3,459 / 10,500 8192 392 / 42 511 / 48 90-96%
K=30 + trim 2,714 / 9,216 6912 446 / 40 707 / 39 52-78%
K=34 + trim 2,375 / 7,168 6144 481 / 40 879 / 49 50-71%
K=36 + trim 2,173 / 6,144 6400 478 / 38 909 / 40 45-60%
K=40 + trim 1,892 / 4,096 5632 499 / 39 800 / 37 34-67%
K=46 + trim 1,537 / 1,024 2816 327 / 34 427 / 34 25-46%

What I take from it (one rig, three runs, please weigh accordingly):

  1. Does the speed term matter on a mixed pair? Yes, and 0.1.39's choice is good. It puts the 3070 first and gives it 43 of 48 layers; that is 1.4-1.8x the P40's prefill and 35-42 decode. The best forced split I found (K=34-36 with the trim) reads a 32K prompt about 45% faster than auto (879-909 vs 617), at the same decode. The opposite order is the bad case: with the P40 first the chunk falls to 1024 (the 3070's room sets it, as @gopinath87607 described) and the pair is slower than the P40 alone.
  2. Prefill follows how many layers the faster card runs, not the hit rate. The hit rate falls from 98% (K=8) to 45-60% (K=34-36) while 32K prefill goes 362 -> 909. Decode is nearly flat (34-42) across all of it, except with the carve.
  3. The carve is the only change that moved decode at an unchanged split (K=20: 38 -> 42 / 48 tok/s, hit rate 82-89% -> 90-96%, chunk 5120 -> 8192). It is opt-in and needs explicit split points, so auto cannot use it. Would you take a change that lets --layer-split auto use it (the cost model already counts what trimming frees)? That is where I would look first.
  4. A worry about the cost model: for auto it printed "the caches hold 4146 of 24576 pairs (~91.7% of the routed mass)", while the decode hit rates measured 39-58%. On IQ2_XS auto chose K=47 (3070 holds 47 layers, 512 slots on the P40), prefill 175-189 and decode 26-27: slower than the P40 alone. With K=36 + trim the same quant gave 527 / 990 prefill. This looks like the mass-curve problem gopinath87607 describes.
  5. --stats: I could not combine it with a split: --stats prints only in one-shot mode and --layer-split needs --serve. For the single P40 I have the breakdown (above, earlier comment); for splits I only have the engine log (slots, chunk, hit rate, the cost model's per-card ms). If there is a flag or /metrics field that gives per-stage time on a split in serve mode, tell me and I will rerun.
  6. Not helpful here, for the record: --peer-device (no P2P between these cards: the log says the prompt path stays on the primary; decode stayed 32-42), --kv int8, --spec 2/3/6, STRATA_BF16_TC=1 (all within noise), --prefill below auto's choice (-4% to -24%).

Still to do on this rig: IQ3_S and Q2_0, concurrency (--batch), Hardin22's fork on this pair, a 32 GB-RAM boot, and 2 x P40 (one card on the chipset x4 slot). I will send those as separate notes. Logs, configs and the runner scripts are available if useful.

Attachments: the engine logs (strata-<model>.log) for each row, the generated configs, and the runner scripts. Related: #584 (placement search, mass curve), #639 / #559 (weight carve), #583 (ring), #642 (resident mode on a split).

Lingua principale
C++
Stelle
11.6k
Fork
1k
Merge medio
7h 46m
PR unite (30g)
30

Preparare l'ambiente

Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di Niko1221/Strata

Tutte le issue di Niko1221/Strata

Issue simili

Altre issue su C++

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.