not stable start with 3+ cards
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 48/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Active
- Domain
- ai-infra-agents, performance
Research direction
Start with the engine log excerpt in the issue and the startup path in serve/server.py, especially the reported failure to allocate the expert cache on CUDA2. Compare the logged free VRAM and cache allocations across the GPUs, then look for the engine code that reports or handles those allocations. Reproduce the startup failure with the reported three RTX 3060 12 GB setup; done means the engine starts reliably with three or more cards without the cache allocation error.
Written by the indexing model from the issue text.
Description
hello! about 10-15 starts and only 2 sucsees. Getting error out of memory for 1 2 or all cards.
try: fixex expert cashe, placing cache at 3rd card, spliting layers like 16,32.
able to start only with 2 cards option w/o vision and not stable.
3060-12gb X3
json file auto created with 2 cars choose have strata-swift-iq2_xs.json "gpu": [ 0, 1, 2 ]
re running setup causing sometimes auto filling options like 1 1 2 2 1 e.t.c without keys press
[strata] experts loaded: 33.02 GiB at 0.07 GiB/s (579 s so far)
[strata] filling the GPU's expert cache (4884 experts, 6.33 GiB of VRAM) ...
[strata] still starting (599 s) - please wait ...
Traceback (most recent call last):
File "F:!strada\serve\server.py", line 4145, in
sys.exit(main())
~~~~^^
File "F:!strada\serve\server.py", line 4001, in main
engine = StrataEngine(exe, engine_args(cfg) + (effort_end or []), cwd=cfg.get("cwd"), log=cfg.get("log"),
env=env, lazy=lazy)
File "F:!strada\serve\server.py", line 478, in init
raise RuntimeError("the engine exited before it was ready" + (f" (see {log})" if log else "") +
start_failure_hint(log, log_start) + start_log_tail(log, log_start))
RuntimeError: the engine exited before it was ready (see F:!strada\strata-swift-iq2_xs.log)
the engine log's last lines:
strata generate: PLE on, table 320001536 rows of F:!strada\Strata-data\models\swift-IQ2_XS\Swift-Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS-00001-of-00002.gguf
strata generate: layer split: CUDA1 holds its weights, session [16, 33); 7.79 GiB free
strata generate: layer split: CUDA2 holds its weights, session [33, 48) and the head; 7.47 GiB free
strata mtp: draft layer loaded, 839 MiB of VRAM (experts 675, dense 111), files read in 9.59 s (82 MiB/s)
strata generate: GPU 0: NVIDIA GeForce RTX 3060, compute capability 8.6
strata generate: multi-GPU under WDDM: at most 8 GiB of the expert arena is pinned (STRATA_ARENA_PIN_GIB changes it)
strata generate: expert arena read unbuffered (0 of 32 probe reads from the file cache; 16.8 GiB available, 63.5 GiB of files)
strata generate: expert arena: locked 25626 MiB via working-set minimum + VirtualLock; cudaHostRegister limited to 8 GiB by the engine (multi-GPU under WDDM, or remote experts); 12 slices pinned (7 GiB); large pages refused for 35456548864 B (GetLargePageMinimum=2097152, VirtualAlloc error 1314); using 4 KB pages
strata generate: loaded 33.02 GiB at 0.07 GiB/s
strata generate: hint: ~24x below what this hardware streams from a normal launch. If Strata is started by Task Scheduler or a service, register the task with Priority 4 (Normal) and 'Run with highest privileges' - the scheduler's defaults (Below normal + a least-privilege token) throttle the load. See docs/DETAILS.md ('Running it at startup').
strata generate: expert cache 4884 slots, 6.33 GiB of VRAM; policy is
strata generate: the GPU computes the experts in the cache; it rounds differently from the CPU,
so a reply can differ slightly from a run without the cache (same quality:
bench/results/2026-09-27-cache-parity).
PROFILE, ranked by routing frequency, no eviction.
strata generate: pre-filled 4884 of 4884 slots from the profile; slot 0 verified
strata generate: layer split, CUDA1: 7.79 GiB free of 12.00, room for experts 7.01 GiB
strata generate: layer split: CUDA1 runs layers 16-32, expert cache 5198 slots (7.01 GiB), 5198 of its 8704 profiled pairs; slot 0 verified
strata generate: layer split, CUDA2: 6.65 GiB free of 12.00, room for experts 5.87 GiB
strata generate: layer split, CUDA2 expert cache: ExpertCache: cudaMalloc(5.87 GiB) failed: out of memory
- Dominant language
- C++
- Stars
- 11.6k
- Forks
- 1k
- Avg merge
- 7h 46m
- Merged PRs (30d)
- 30
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Niko1221/Strata
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Niko1221/Strata#974 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
Similar issues
-
area:runtime good first issue
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
WATonomous/wato_f1tenth#39 ·
-
[APP BUG]: Sorting by name after searching can bring up irrelevant resultsPossibly taken A pull request linked to this issue is open or already merged. Open
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
shadps4-emu/shadps4-qtlauncher#465 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
duckdb/duckdb-excel#104 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
lxqt/lxqt-powermanagement#495 ·
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Qiskit/qiskit-aer#2466 ·