P40 + 3070 (mixed Pascal/Ampere pair): `layer_split: auto` vs forced splits on 0.1.39, with hit rates (follow-up to #604)
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 35/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- cpp
- Domain
- backend, performance
Research direction
Start with the related placement-search and weight-carve issues (#584, #639, and #559), then inspect how the layer-split cost model and auto selection account for trimming. The issue names bench/results/.../strata-bench.py and provides configs and engine logs for comparison. Done means determining whether auto can use trimming and validating the resulting split and performance against the reported forced-split runs.
Written by the indexing model from the issue text.
Description
Follow-up to #604, where you asked for "a P40 + 3070 with layer_split: auto against a few forced splits (--stats and the hit rates)". Here is that run, on v0.1.39 as released (no patches; CUDA 12.9 engine, STRATA_EXPERIMENTAL_SM60=1), IQ3_XXS, 36,864-token context, P40 power-capped at 130 W, DDR4-2400, PCIe 3.0 x8. Three runs per cell with bench/results/.../strata-bench.py (unique prefix per run, 256 tokens). Prefill / decode in tok/s; hit = decode expert-cache hit rate range over the runs.
| config (3070 first unless noted) | slots 3070 / P40 | chunk | 4K | 32K | hit |
|---|---|---|---|---|---|
| P40 alone | - / 10,451 | 8192 | 322 / 34 | 339 / 34 | 83-96% |
| 3070 alone | 513 / - | 1024 | 144 / 30 | 155 / 37 | 14-29% |
| auto (K=43) | 1,536 / 2,560 | 4096 | 462 / 35 | 617 / 42 | 39-58% |
| auto, P40 first | 8,192 / 736 | 1024 | 193 / 32 | 221 / 37 | 26-42% |
| forced K=8 | 2,177 / 10,333 | 5888 | 333 / 39 | 362 / 37 | 89-98% |
| forced K=20 | 2,070 / 9,780 | 5120 | 386 / 38 | 491 / 38 | 82-89% |
K=20 + STRATA_STAGE_TRIM=1 |
3,459 / 10,500 | 8192 | 392 / 42 | 511 / 48 | 90-96% |
| K=30 + trim | 2,714 / 9,216 | 6912 | 446 / 40 | 707 / 39 | 52-78% |
| K=34 + trim | 2,375 / 7,168 | 6144 | 481 / 40 | 879 / 49 | 50-71% |
| K=36 + trim | 2,173 / 6,144 | 6400 | 478 / 38 | 909 / 40 | 45-60% |
| K=40 + trim | 1,892 / 4,096 | 5632 | 499 / 39 | 800 / 37 | 34-67% |
| K=46 + trim | 1,537 / 1,024 | 2816 | 327 / 34 | 427 / 34 | 25-46% |
What I take from it (one rig, three runs, please weigh accordingly):
- Does the speed term matter on a mixed pair? Yes, and 0.1.39's choice is good. It puts the 3070 first and gives it 43 of 48 layers; that is 1.4-1.8x the P40's prefill and 35-42 decode. The best forced split I found (K=34-36 with the trim) reads a 32K prompt about 45% faster than auto (879-909 vs 617), at the same decode. The opposite order is the bad case: with the P40 first the chunk falls to 1024 (the 3070's room sets it, as @gopinath87607 described) and the pair is slower than the P40 alone.
- Prefill follows how many layers the faster card runs, not the hit rate. The hit rate falls from 98% (K=8) to 45-60% (K=34-36) while 32K prefill goes 362 -> 909. Decode is nearly flat (34-42) across all of it, except with the carve.
- The carve is the only change that moved decode at an unchanged split (K=20: 38 -> 42 / 48 tok/s, hit rate 82-89% -> 90-96%, chunk 5120 -> 8192). It is opt-in and needs explicit split points, so
autocannot use it. Would you take a change that lets--layer-split autouse it (the cost model already counts what trimming frees)? That is where I would look first. - A worry about the cost model: for auto it printed "the caches hold 4146 of 24576 pairs (~91.7% of the routed mass)", while the decode hit rates measured 39-58%. On IQ2_XS auto chose K=47 (3070 holds 47 layers, 512 slots on the P40), prefill 175-189 and decode 26-27: slower than the P40 alone. With K=36 + trim the same quant gave 527 / 990 prefill. This looks like the mass-curve problem gopinath87607 describes.
--stats: I could not combine it with a split:--statsprints only in one-shot mode and--layer-splitneeds--serve. For the single P40 I have the breakdown (above, earlier comment); for splits I only have the engine log (slots, chunk, hit rate, the cost model's per-card ms). If there is a flag or/metricsfield that gives per-stage time on a split in serve mode, tell me and I will rerun.- Not helpful here, for the record:
--peer-device(no P2P between these cards: the log says the prompt path stays on the primary; decode stayed 32-42),--kv int8,--spec 2/3/6,STRATA_BF16_TC=1(all within noise),--prefillbelow auto's choice (-4% to -24%).
Still to do on this rig: IQ3_S and Q2_0, concurrency (--batch), Hardin22's fork on this pair, a 32 GB-RAM boot, and 2 x P40 (one card on the chipset x4 slot). I will send those as separate notes. Logs, configs and the runner scripts are available if useful.
Attachments: the engine logs (strata-<model>.log) for each row, the generated configs, and the runner scripts. Related: #584 (placement search, mass curve), #639 / #559 (weight carve), #583 (ring), #642 (resident mode on a split).
- Dominant language
- C++
- Stars
- 11.6k
- Forks
- 1k
- Avg merge
- 7h 46m
- Merged PRs (30d)
- 30
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Niko1221/Strata
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Niko1221/Strata#974 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
shadps4-emu/shadps4-qtlauncher#465 ·
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Qiskit/qiskit-aer#2466 ·
-
bug needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
NVIDIA/attestation-sdk#41 · 1 comment ·
-
[Bug]: CMAKE Fails to find libgit2 on POP OSPossibly taken @Tirpitz93 claimed this today. Openbug
Difficulty 1/5 Under an hour Newbie friendliness 85/100
subsurface/subsurface#5007 · 2 comments ·
Maintainers usually reply within 2 days
-
feature request
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 2 days