The profile's length caps the expert arena — even an explicit --expert-cache N
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 35/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Domain
- backend, performance
Research direction
Read the cache sizing and admission paths in src/program/generate.cpp, then review the policy question in include/strata/core/expert_cache.hpp:107 and the profile requirement in src/core/verify.cpp:325. First determine whether profiles are intended as a hard capacity ceiling or as seeds for on-demand admission; the issue does not settle that design choice. Done means the intended behavior is decided and documented, with explicit truncation no longer silent if the ceiling is intentional.
Written by the indexing model from the issue text.
Description
With --expert-profile, the profile's length becomes the ceiling of the VRAM expert arena: auto is capped at it, an explicit --expert-cache N above it is silently truncated to it (native pack), and the one path that keeps N slots (--expert-cache-per-layer) can never fill the ones past the profile. Measurements below. Is the profile's length meant to be the ceiling? Or should the profile seed the hot experts and the remaining slots fill on demand - which is what a user with a trace-built profile (make_profile.py --no-base) and a big card most likely expects?
The three mechanisms:
--expert-cache autois capped atprofile.size()(src/program/generate.cpp:3222). Known and documented in tools/make_profile.py (issue #46 was fixed by shipping profiles that rank every pair) - fine.- An explicit
--expert-cache Nlarger than the profile is silently truncated to the profile's length on the native-pack path: the sized-slots loop only iterates the profile's pairs and rewriteso.expert_cache(src/program/generate.cpp:3299-3314). No warning is printed. --expert-cache-per-layerskips that clamp, so the arena keeps N slots - but the profile is the policy ("the decode-time admission finds no room and every non-profiled expert stays a CPU miss", generate.cpp:3451-3454), so the slots past the profile can never hold anything. The VRAM is allocated and permanently idle.
And since --serve requires the profile-filled tier (Verifier::init, src/core/verify.cpp:325 - the engine exits at startup without --expert-profile), the compulsory-miss fill the engine already implements for profile-less runs (generate.cpp:3445) is unreachable exactly where it would be useful: serving. This all seems adjacent to R4.1's open admission/eviction question (include/strata/core/expert_cache.hpp:107).
Measured
RTX 3080 20 GB (driver 595.91.07, CUDA 13.2), Xeon E5-2680 v4, 62 GB RAM, Ubuntu 24.04.5, engine 0.1.39 built from source, Qwen3.8-Flash-Next GSQ-RCO IQ2_XS (48 x 512 = 24,576 pairs), --max-context 262144 --kv int8 --kv-resident 32768 --spec 4 --prefill auto. Same model and flags throughout; only the profile and --expert-cache vary.
A. Full profile (24,576 pairs), --expert-cache auto - the shipped behavior, fine:
strata generate: expert cache 9813 slots, 11.66 GiB of VRAM; policy is
PROFILE, ranked by routing frequency, no eviction.
B. Profile truncated to 3,000 pairs, explicit --expert-cache 9813 - asked for 9,813, got 3,000, no warning:
strata generate: profile .../expert-profile-3000.bin: 3000 ranked pairs, built for 3000 slots
strata generate: expert cache 3000 slots, 4.04 GiB of VRAM; policy is
PROFILE, ranked by routing frequency, no eviction.
strata generate: pre-filled 3000 of 3000 slots from the profile; slot 0 verified
C. Same profile, --expert-cache-per-layer, explicit --expert-cache 8000 - the arena keeps 8,000 slots (11.25 GiB), 3,000 are filled, the remaining 5,000 (~7 GiB) can never hold an expert:
strata generate: expert cache 8000 slots, 11.25 GiB of VRAM; policy is
R4.2g PER-LAYER: each layer owns 166 slots (0..165).
strata generate: pre-filled 3000 of 8000 slots from the profile; slot 0 verified
(The 3,000-of-8,000 residency was cross-checked with a local patch that reports ExpertCache::resident() in the engine's INFO/DONE lines and the Monitor, where the "experts cached" figure currently reads the capacity instead. Happy to send that as a small PR if you want it.)
Reproducing
Truncate any profile to its first 3,000 pairs (the format is make_profile.py's):
import sys; sys.path.insert(0, ".")
from tools.make_profile import read_profile, write_profile
write_profile("data/expert-profile-3000.bin", read_profile("data/expert-profile.bin")[:3000])
Then start the server with --expert-profile data/expert-profile-3000.bin and either --expert-cache 9813 (case B) or --expert-cache 8000 --expert-cache-per-layer (case C), and read the startup log.
What I'd suggest
- If the ceiling is intended: say so in
--help/ docs/DETAILS.md, and warn when an explicit N is truncated (case B is silent today). - If not: the smallest step is probably letting the slots past the profile fill on compulsory misses - the admission path exists, it is the verify tier's profile requirement that keeps it out of
--serve. I can measure the hit-rate difference (static profile vs profile-seeded hybrid) on the 3080 if that helps the R4.1 question.
Drafted with AI assistance (Claude); every number above was measured on the machine listed, and every file:line reference was checked against the source at 0.1.39.
- Dominant language
- C++
- Stars
- 11.6k
- Forks
- 1k
- Avg merge
- 7h 46m
- Merged PRs (30d)
- 30
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Niko1221/Strata
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Niko1221/Strata#974 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
Similar issues
-
feature request
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 2 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
gavinlouuu-kpt/mib-studio-qt#517 ·
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 73/100
EchoTools/nevr-runtime#116 ·
Maintainers usually reply within 1 day
-
code-quality libc++
Difficulty 1/5 Under an hour Newbie friendliness 82/100
llvm/llvm-project#229284 ·
Maintainers usually reply within 1 day