[Optimization] Support CPPC / configurable core placement for host loop and expert pool
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 80/100
- Issue type
- Feature
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- cpp
- Domain
- performance
Research direction
Start at src/kernels/cpu/pool.cpp in CorePlacement::CorePlacement, where host_ and workers_ are chosen; read how host_ and workers_ are consumed downstream (host thread pinning and worker pool) before editing. Implement the STRATA_HOST_PROC / STRATA_WORKER_PROCS getenv overrides with bounds checks against core_of_, and confirm that with no env vars set the existing placement logic is unchanged. Done means a user can set those variables (including via the env block in strata-<model>.json) to pin the host loop to a chosen logical processor and supply a comma-separated worker list, with out-of-range values ignored rather than crashing.
Written by the indexing model from the issue text.
Description
Feature / Optimization: Support CPPC preferred core placement for host loop and expert pool (avoid pinning host loop to slowest core)
Context
In src/kernels/cpu/pool.cpp, CorePlacement::CorePlacement initializes the placement for the host loop and the worker pool:
CorePlacement::CorePlacement(bool spare_first) : cores_(physical_cores()) {
for (size_t c = 0; c < cores_.size(); ++c)
for (int p : cores_[c]) {
if (p >= (int) core_of_.size()) core_of_.resize((size_t) p + 1, -1);
core_of_[(size_t) p] = (int) c;
}
host_ = cores_.back().front();
for (size_t c = spare_first && cores_.size() > 4 ? 1 : 0; c + 1 < cores_.size(); ++c)
workers_.push_back(cores_[c].front());
}
Problem
On modern multi-core x86 processors with asymmetric core performance (AMD Zen 3/4/5 ACPI CPPC preferred cores, Intel hybrid P/E architectures), silicon binning varies across physical cores on the die.
Hardcoding host_ = cores_.back().front() blindly pins the latency-critical host thread to the highest-numbered physical core (e.g. Core 7 on an 8-core CPU).
On an AMD Ryzen 7 5800X3D (8C/16T, dual RTX 3090 setup), querying the Windows kernel ACPI CPPC power capabilities (Kernel-Processor-Power Event ID 55) reveals:
- Core 2 (Processors 4, 5): 158% max performance capability (CPPC Preferred / Gold Star)
- Core 3 (Processors 6, 7): 158% max performance capability (CPPC Preferred / Gold Star)
- Core 1 (Processors 2, 3): 154%
- Core 4 (Processors 8, 9): 150%
- Core 0 (Processors 0, 1): 145% (intentionally spared for GPU PCIe interrupts and OS DPCs)
- Core 5 (Processors 10, 11): 141%
- Core 6 (Processors 12, 13): 137%
- Core 7 (Processors 14, 15): 133% (the slowest, lowest-binned core on the entire chip)
Because the host thread runs session_loop, spins on cudaEventQuery, schedules CUDA kernels across GPUs, conducts MTP verification, and handles token sampling, pinning it to the slowest silicon core caps its single-threaded boost frequency and introduces unnecessary scheduling latency.
Proposed Solution
- Configurable Environment Variables:
Allow runtime overrides via environment variables (STRATA_HOST_PROCandSTRATA_WORKER_PROCS) inCorePlacement::CorePlacement:
const char* env_host = std::getenv("STRATA_HOST_PROC");
if (env_host != nullptr && *env_host != '\0') {
int hp = std::atoi(env_host);
if (hp >= 0 && hp < (int) core_of_.size()) host_ = hp;
}
const char* env_workers = std::getenv("STRATA_WORKER_PROCS");
if (env_workers != nullptr && *env_workers != '\0') {
std::vector<int> w;
std::stringstream ss(env_workers);
std::string item;
while (std::getline(ss, item, ',')) {
int p = std::atoi(item.c_str());
if (p >= 0 && p < (int) core_of_.size()) w.push_back(p);
}
if (!w.empty()) workers_ = w;
}
This allows setting in strata-<model>.json:
"env": {
"STRATA_HOST_PROC": "4",
"STRATA_WORKER_PROCS": "6,2,8,10,12,14"
}
- Automated CPPC Detection (Optional Future Extension):
Alternatively, query OS CPPC information at startup:
- Windows:
GetSystemCpuSetInformation/ ACPI_CPCobjects - Linux:
/sys/devices/system/cpu/cpu*/acpi_cppc/highest_perf
and sort cores by performance before allocatinghost_andworkers_.
Verification & Benchmarks
Testing with Qwen3.8-Flash-Next on Dual RTX 3090s and Ryzen 7 5800X3D:
- By sparing Core 0 (0,1) for GPU interrupts, pinning the Host Thread to Core 2 (processor 4, 158% Gold Star), and assigning workers to Cores 3, 1, 4, 5, 6, 7 (
6,2,8,10,12,14):- Short prompt prefill improved by +6.0% (58.7 tok/s -> 62.2 tok/s).
- Host event loop and CUDA synchronization runs on the highest-clocked core instead of the silicon's worst core.
- Dominant language
- C++
- Stars
- 11.6k
- Forks
- 1k
- Avg merge
- 7h 46m
- Merged PRs (30d)
- 30
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Niko1221/Strata
-
Difficulty 1/5 Under an hour Newbie friendliness 82/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 64/100
Niko1221/Strata#1248 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 1 day
Similar issues
-
Component: Ruby Type: bug
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
Maintainers usually reply within 1 day
-
bug needs triage tcp
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
project-chip/connectedhomeip#74644 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
godotengine/godot#124252 · 2 comments ·
Maintainers usually reply within 1 day
-
`controller_manager/activity` doesn't update after handling hardware errorPossibly taken @firuzakhmad claimed this today. Openbug
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
ros-controls/ros2_control#3668 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 66/100
texstudio-org/texstudio#4685 ·
Maintainers usually reply within 1 day