Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[Optimization] Support CPPC / configurable core placement for host loop and expert pool

Open Beginner friendly
#1,254 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
80/100
Issue type
Feature
Clarity
Clearly specified
Activity status
Active
Tech stack
cpp
Domain
performance

Research direction

Start at src/kernels/cpu/pool.cpp in CorePlacement::CorePlacement, where host_ and workers_ are chosen; read how host_ and workers_ are consumed downstream (host thread pinning and worker pool) before editing. Implement the STRATA_HOST_PROC / STRATA_WORKER_PROCS getenv overrides with bounds checks against core_of_, and confirm that with no env vars set the existing placement logic is unchanged. Done means a user can set those variables (including via the env block in strata-<model>.json) to pin the host loop to a chosen logical processor and supply a comma-separated worker list, with out-of-range values ignored rather than crashing.

Written by the indexing model from the issue text.

Description

Feature / Optimization: Support CPPC preferred core placement for host loop and expert pool (avoid pinning host loop to slowest core)
Context

In src/kernels/cpu/pool.cpp, CorePlacement::CorePlacement initializes the placement for the host loop and the worker pool:

CorePlacement::CorePlacement(bool spare_first) : cores_(physical_cores()) {
    for (size_t c = 0; c < cores_.size(); ++c)
        for (int p : cores_[c]) {
            if (p >= (int) core_of_.size()) core_of_.resize((size_t) p + 1, -1);
            core_of_[(size_t) p] = (int) c;
        }
    host_ = cores_.back().front();
    for (size_t c = spare_first && cores_.size() > 4 ? 1 : 0; c + 1 < cores_.size(); ++c)
        workers_.push_back(cores_[c].front());
}
Problem

On modern multi-core x86 processors with asymmetric core performance (AMD Zen 3/4/5 ACPI CPPC preferred cores, Intel hybrid P/E architectures), silicon binning varies across physical cores on the die.

Hardcoding host_ = cores_.back().front() blindly pins the latency-critical host thread to the highest-numbered physical core (e.g. Core 7 on an 8-core CPU).

On an AMD Ryzen 7 5800X3D (8C/16T, dual RTX 3090 setup), querying the Windows kernel ACPI CPPC power capabilities (Kernel-Processor-Power Event ID 55) reveals:

  • Core 2 (Processors 4, 5): 158% max performance capability (CPPC Preferred / Gold Star)
  • Core 3 (Processors 6, 7): 158% max performance capability (CPPC Preferred / Gold Star)
  • Core 1 (Processors 2, 3): 154%
  • Core 4 (Processors 8, 9): 150%
  • Core 0 (Processors 0, 1): 145% (intentionally spared for GPU PCIe interrupts and OS DPCs)
  • Core 5 (Processors 10, 11): 141%
  • Core 6 (Processors 12, 13): 137%
  • Core 7 (Processors 14, 15): 133% (the slowest, lowest-binned core on the entire chip)

Because the host thread runs session_loop, spins on cudaEventQuery, schedules CUDA kernels across GPUs, conducts MTP verification, and handles token sampling, pinning it to the slowest silicon core caps its single-threaded boost frequency and introduces unnecessary scheduling latency.

Proposed Solution
  1. Configurable Environment Variables:
    Allow runtime overrides via environment variables (STRATA_HOST_PROC and STRATA_WORKER_PROCS) in CorePlacement::CorePlacement:
const char* env_host = std::getenv("STRATA_HOST_PROC");
if (env_host != nullptr && *env_host != '\0') {
    int hp = std::atoi(env_host);
    if (hp >= 0 && hp < (int) core_of_.size()) host_ = hp;
}
const char* env_workers = std::getenv("STRATA_WORKER_PROCS");
if (env_workers != nullptr && *env_workers != '\0') {
    std::vector<int> w;
    std::stringstream ss(env_workers);
    std::string item;
    while (std::getline(ss, item, ',')) {
        int p = std::atoi(item.c_str());
        if (p >= 0 && p < (int) core_of_.size()) w.push_back(p);
    }
    if (!w.empty()) workers_ = w;
}

This allows setting in strata-<model>.json:

"env": {
  "STRATA_HOST_PROC": "4",
  "STRATA_WORKER_PROCS": "6,2,8,10,12,14"
}
  1. Automated CPPC Detection (Optional Future Extension):
    Alternatively, query OS CPPC information at startup:
  • Windows: GetSystemCpuSetInformation / ACPI _CPC objects
  • Linux: /sys/devices/system/cpu/cpu*/acpi_cppc/highest_perf
    and sort cores by performance before allocating host_ and workers_.
Verification & Benchmarks

Testing with Qwen3.8-Flash-Next on Dual RTX 3090s and Ryzen 7 5800X3D:

  • By sparing Core 0 (0,1) for GPU interrupts, pinning the Host Thread to Core 2 (processor 4, 158% Gold Star), and assigning workers to Cores 3, 1, 4, 5, 6, 7 (6,2,8,10,12,14):
    • Short prompt prefill improved by +6.0% (58.7 tok/s -> 62.2 tok/s).
    • Host event loop and CUDA synchronization runs on the highest-clocked core instead of the silicon's worst core.
Dominant language
C++
Stars
11.6k
Forks
1k
Avg merge
7h 46m
Merged PRs (30d)
30

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Niko1221/Strata

All issues in Niko1221/Strata

Similar issues

More C++ issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.