test_cpu_threadpool: the 100x dispatch ratio is calibrated on 20 cores and flakes on 4-vCPU CI runners
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 3/5
- 見積もり時間
- 1〜2日
- 初心者へのやさしさ
- 58/100
- issue の種類
- バグ
- 明瞭さ
- おおむね明確
- 活発さ
- 活発
- 技術スタック
- cpp
- 領域
- testing-qa
調査の方向性
tests/vt/test_cpu_threadpool.cpp の oversubscribed dispatch does not cost a scheduler timeslice ケースから始め、Threadpool::Run のそばにあるキャリブレーションコメントを読んでください。4 vCPU の runner で dispatch 比率を再現し、その後、記録された scheduler timeslice の不具合に対する感度を維持するキャリブレーション調整を選択して検証してください。CPU test job を実行し、小規模な runner でそのケースが成功することを確認するとともに、その不具合を隠していないことを確認してください。
索引モデルが issue の本文から書いたものです。
説明
Row: QUANT-GGUF-CPU-THREADPOOL
tests/vt/test_cpu_threadpool.cpp, case oversubscribed dispatch does not cost a scheduler timeslice, asserts ratio < 100.0 on ratio = median_dispatch_us(cores + 1) / median_dispatch_us(cores / 2).
Observed on CI (4-vCPU hosted runner)
build-test-cpu job 102683560829 (PR #3095, run 34416865779), tests/vt/test_cpu_threadpool.cpp:610:
MESSAGE: empty-op dispatch: 2 threads 0.421 us, 5 threads 42.75 us, ratio 101.544
ERROR: CHECK( ratio < 100.0 ) is NOT correct!
The arms are 2 and 5 threads, which proves the runner reported hardware_concurrency() == 4.
Why this reads as a calibration gap rather than a detected defect
- The recorded defect signature in this test's own comment is an over arm of 2999-5996 us at 21 threads against 18-20 us fixed. Here the over arm is 42.75 us, roughly 70x below that signature, so the waiter was not spinning through a scheduler timeslice.
- The comment's calibration ("Verified both ways on 20 cores") and its margin ("~30x below the defect and ~28x above the fixed behaviour") come from a 20-core box with arms 10 and 21. On a 4-vCPU runner the arms are 2 and 5 and the denominator measured 0.421 us, about 4x below the 1.68 us the same comment calibrates against, so the ratio is dominated by denominator noise.
std::thread::hardware_concurrency()is not reduced by CPU affinity on this glibc:taskset -c 0-3 <binary>still selects 24/49-thread arms, so the 4-vCPU geometry cannot be reproduced locally by pinning cores.main's ownbuild-test-cpujob103074166342passes the same test on the same runner class, so the check is marginal there rather than deterministically red.
Suggested fix (owner's call; changing an assertion needs its own spec)
Floor the denominator, for example over_us < std::max(100.0 * fits_us, <absolute floor>), or scale the threshold with cores, or skip the case below 8 cores. Any of these must keep the case able to fail for the defect it was written for.
Found while triaging an unrelated red check on PR #3095: that diff changes attention and embedding gating in src/vt/ops.cpp and never calls the measured path, which is Threadpool::Run with an empty body.
- 主要言語
- C++
- スター
- 423
- フォーク
- 53
- 平均マージ
- 1日 7時間
- マージ済み PR(30日)
- 382
環境構築
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
mudler/vllm.cpp のほかの issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
メンテナーはふだん 1 日以内に返信
mudler/vllm.cpp の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
hyprwm/aquamarine#426 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 62/100
amnezia-vpn/amnezia-client#3222 ·
メンテナーはふだん 2 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
sdatkinson/NeuralAmpModelerCore#342 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
valkey-io/valkey-search#1465 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
KhronosGroup/Vulkan-Tutorial#524 ·
メンテナーはふだん 1 日以内に返信