docs: MULTI_GPU.md lists Pascal as unsupported for the layer split, but 2x Tesla P40 runs it (v0.1.39)
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 68/100
- Issue type
- Documentation
- Clarity
- Mostly clear
- Activity status
- Active
- Domain
- documentation
Research direction
Start with docs/MULTI_GPU.md and docs/OLDER_GPUS.md, then compare their Pascal and multi-GPU guidance with the behavior reported in this issue and the results in #1028. Clarify whether two Pascal cards are experimentally supported and whether manual config editing is required; update the relevant documentation so the two pages agree.
Written by the indexing model from the issue text.
Description
Observation (v0.1.39, main = 6f32ec0, unmodified): two Tesla P40 22 GB (compute capability 6.1) ran Flash-Next IQ2_XS with the layer split, although docs/MULTI_GPU.md lists cards below compute capability 7.5 as unsupported.
How: I installed with ./setup.sh ... --cuda 12 --build --gpu 1 (CUDA 12.4 + g++-13, engine compiled for sm_61 with STRATA_EXPERIMENTAL_SM60=1), then edited the generated config by hand to "gpu": [0, 1] and "layer_split": "auto" and started serve/server.py --engine strata --config .... I did not try --gpus 0,1 through setup, so I can't say whether setup itself would refuse it.
Evidence (engine log, two-card runs):
strata generate: layer split across 2 GPUs: CUDA0, then CUDA1 (split auto)
strata generate: layer split auto: K=24 - predicted 75.8 ms per decode window; the caches hold 24576 of 24576 profiled pairs (~100.0% of the routed mass)
strata generate: layer split: CUDA1 holds its weights, session [24, 48) and the head; 18.63 GiB free
Median decode for a short prompt went from 19.6 tok/s (one P40) to 34.8 tok/s (both); prompt reading 348 / 388 tok/s (4K / 16K) on one card and 353 / 501 on both. Three runs per cell; the first run of each is much slower (warm-up). All per-run data, engine logs and the config files are in #1028 (a results-only community report).
Docs: docs/MULTI_GPU.md ("Not supported", and the example listing a GTX 1080 Ti as "not supported - older than the RTX 20 series") and docs/OLDER_GPUS.md (which allows --gpus 0,1 "with a newer card" and says a V100 can share a model with an RTX 30/40 card) read differently for two Pascal cards, which neither page mentions. If this is expected to work experimentally, a sentence in either page would help; if it only works by hand-editing the config, that may be worth saying too.
Caveats: one machine, one quantisation (IQ2_XS), 32,768 context, three runs; the two-card runs were not NUMA-pinned. Not tested: --batch with a layer split, vision, contexts above 32K, other quantisations. Happy to test a specific variant if that helps (for example --gpus 0,1 --cuda 12 through setup, or IQ3_S).
Related Pascal reports: #875, #395. The wider study (llama.cpp and Ollama on the same weights, accuracy and repeatability): https://github.com/sarge18/p40-llm-engine-bakeoff
Disclosure: prepared with an AI assistant (Claude Sonnet 5.5, medium effort) under my direction; every number above comes from the per-run data in #1028.
- Dominant language
- C++
- Stars
- 11.6k
- Forks
- 1k
- Avg merge
- 7h 46m
- Merged PRs (30d)
- 30
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Niko1221/Strata
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Niko1221/Strata#974 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
Similar issues
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 73/100
EchoTools/nevr-runtime#116 ·
Maintainers usually reply within 1 day
-
code-quality libc++
Difficulty 1/5 Under an hour Newbie friendliness 82/100
llvm/llvm-project#229284 ·
Maintainers usually reply within 1 day
-
test-issue
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
llvm/offload-test-suite#1557 ·
Maintainers usually reply within 1 day
-
enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
iOS: hidden scale bar invalidates its intrinsic content size on every layout pass of MLNMapViewOpen
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
maplibre/maplibre-native#4723 ·
Maintainers usually reply within 1 day