Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

P40 + 3070 (mixed Pascal/Ampere pair): `layer_split: auto` vs forced splits on 0.1.39, with hit rates (follow-up to #604)

Open
#875 2 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
35/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Active
Tech stack
cpp

Research direction

Start with the related placement-search and weight-carve issues (#584, #639, and #559), then inspect how the layer-split cost model and auto selection account for trimming. The issue names bench/results/.../strata-bench.py and provides configs and engine logs for comparison. Done means determining whether auto can use trimming and validating the resulting split and performance against the reported forced-split runs.

Written by the indexing model from the issue text.

Description

p40-3070-logs-v0.1.39.tar.gz

Follow-up to #604, where you asked for "a P40 + 3070 with layer_split: auto against a few forced splits (--stats and the hit rates)". Here is that run, on v0.1.39 as released (no patches; CUDA 12.9 engine, STRATA_EXPERIMENTAL_SM60=1), IQ3_XXS, 36,864-token context, P40 power-capped at 130 W, DDR4-2400, PCIe 3.0 x8. Three runs per cell with bench/results/.../strata-bench.py (unique prefix per run, 256 tokens). Prefill / decode in tok/s; hit = decode expert-cache hit rate range over the runs.

config (3070 first unless noted) slots 3070 / P40 chunk 4K 32K hit
P40 alone - / 10,451 8192 322 / 34 339 / 34 83-96%
3070 alone 513 / - 1024 144 / 30 155 / 37 14-29%
auto (K=43) 1,536 / 2,560 4096 462 / 35 617 / 42 39-58%
auto, P40 first 8,192 / 736 1024 193 / 32 221 / 37 26-42%
forced K=8 2,177 / 10,333 5888 333 / 39 362 / 37 89-98%
forced K=20 2,070 / 9,780 5120 386 / 38 491 / 38 82-89%
K=20 + STRATA_STAGE_TRIM=1 3,459 / 10,500 8192 392 / 42 511 / 48 90-96%
K=30 + trim 2,714 / 9,216 6912 446 / 40 707 / 39 52-78%
K=34 + trim 2,375 / 7,168 6144 481 / 40 879 / 49 50-71%
K=36 + trim 2,173 / 6,144 6400 478 / 38 909 / 40 45-60%
K=40 + trim 1,892 / 4,096 5632 499 / 39 800 / 37 34-67%
K=46 + trim 1,537 / 1,024 2816 327 / 34 427 / 34 25-46%

What I take from it (one rig, three runs, please weigh accordingly):

  1. Does the speed term matter on a mixed pair? Yes, and 0.1.39's choice is good. It puts the 3070 first and gives it 43 of 48 layers; that is 1.4-1.8x the P40's prefill and 35-42 decode. The best forced split I found (K=34-36 with the trim) reads a 32K prompt about 45% faster than auto (879-909 vs 617), at the same decode. The opposite order is the bad case: with the P40 first the chunk falls to 1024 (the 3070's room sets it, as @gopinath87607 described) and the pair is slower than the P40 alone.
  2. Prefill follows how many layers the faster card runs, not the hit rate. The hit rate falls from 98% (K=8) to 45-60% (K=34-36) while 32K prefill goes 362 -> 909. Decode is nearly flat (34-42) across all of it, except with the carve.
  3. The carve is the only change that moved decode at an unchanged split (K=20: 38 -> 42 / 48 tok/s, hit rate 82-89% -> 90-96%, chunk 5120 -> 8192). It is opt-in and needs explicit split points, so auto cannot use it. Would you take a change that lets --layer-split auto use it (the cost model already counts what trimming frees)? That is where I would look first.
  4. A worry about the cost model: for auto it printed "the caches hold 4146 of 24576 pairs (~91.7% of the routed mass)", while the decode hit rates measured 39-58%. On IQ2_XS auto chose K=47 (3070 holds 47 layers, 512 slots on the P40), prefill 175-189 and decode 26-27: slower than the P40 alone. With K=36 + trim the same quant gave 527 / 990 prefill. This looks like the mass-curve problem gopinath87607 describes.
  5. --stats: I could not combine it with a split: --stats prints only in one-shot mode and --layer-split needs --serve. For the single P40 I have the breakdown (above, earlier comment); for splits I only have the engine log (slots, chunk, hit rate, the cost model's per-card ms). If there is a flag or /metrics field that gives per-stage time on a split in serve mode, tell me and I will rerun.
  6. Not helpful here, for the record: --peer-device (no P2P between these cards: the log says the prompt path stays on the primary; decode stayed 32-42), --kv int8, --spec 2/3/6, STRATA_BF16_TC=1 (all within noise), --prefill below auto's choice (-4% to -24%).

Still to do on this rig: IQ3_S and Q2_0, concurrency (--batch), Hardin22's fork on this pair, a 32 GB-RAM boot, and 2 x P40 (one card on the chipset x4 slot). I will send those as separate notes. Logs, configs and the runner scripts are available if useful.

Attachments: the engine logs (strata-<model>.log) for each row, the generated configs, and the runner scripts. Related: #584 (placement search, mass curve), #639 / #559 (weight carve), #583 (ring), #642 (resident mode on a split).

Dominant language
C++
Stars
11.6k
Forks
1k
Avg merge
7h 46m
Merged PRs (30d)
30

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Niko1221/Strata

All issues in Niko1221/Strata

Similar issues

More C++ issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.