Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

test_cpu_threadpool: the 100x dispatch ratio is calibrated on 20 cores and flakes on 4-vCPU CI runners

クローズ
#3,130 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
3/5
見積もり時間
1〜2日
初心者へのやさしさ
58/100
issue の種類
バグ
明瞭さ
おおむね明確
活発さ
活発
技術スタック
cpp
領域
testing-qa

調査の方向性

tests/vt/test_cpu_threadpool.cpp の oversubscribed dispatch does not cost a scheduler timeslice ケースから始め、Threadpool::Run のそばにあるキャリブレーションコメントを読んでください。4 vCPU の runner で dispatch 比率を再現し、その後、記録された scheduler timeslice の不具合に対する感度を維持するキャリブレーション調整を選択して検証してください。CPU test job を実行し、小規模な runner でそのケースが成功することを確認するとともに、その不具合を隠していないことを確認してください。

索引モデルが issue の本文から書いたものです。

説明

Row: QUANT-GGUF-CPU-THREADPOOL

tests/vt/test_cpu_threadpool.cpp, case oversubscribed dispatch does not cost a scheduler timeslice, asserts ratio < 100.0 on ratio = median_dispatch_us(cores + 1) / median_dispatch_us(cores / 2).

Observed on CI (4-vCPU hosted runner)

build-test-cpu job 102683560829 (PR #3095, run 34416865779), tests/vt/test_cpu_threadpool.cpp:610:

MESSAGE: empty-op dispatch: 2 threads 0.421 us, 5 threads 42.75 us, ratio 101.544
ERROR: CHECK( ratio < 100.0 ) is NOT correct!

The arms are 2 and 5 threads, which proves the runner reported hardware_concurrency() == 4.

Why this reads as a calibration gap rather than a detected defect

  • The recorded defect signature in this test's own comment is an over arm of 2999-5996 us at 21 threads against 18-20 us fixed. Here the over arm is 42.75 us, roughly 70x below that signature, so the waiter was not spinning through a scheduler timeslice.
  • The comment's calibration ("Verified both ways on 20 cores") and its margin ("~30x below the defect and ~28x above the fixed behaviour") come from a 20-core box with arms 10 and 21. On a 4-vCPU runner the arms are 2 and 5 and the denominator measured 0.421 us, about 4x below the 1.68 us the same comment calibrates against, so the ratio is dominated by denominator noise.
  • std::thread::hardware_concurrency() is not reduced by CPU affinity on this glibc: taskset -c 0-3 <binary> still selects 24/49-thread arms, so the 4-vCPU geometry cannot be reproduced locally by pinning cores.
  • main's own build-test-cpu job 103074166342 passes the same test on the same runner class, so the check is marginal there rather than deterministically red.

Suggested fix (owner's call; changing an assertion needs its own spec)

Floor the denominator, for example over_us < std::max(100.0 * fits_us, <absolute floor>), or scale the threshold with cores, or skip the case below 8 cores. Any of these must keep the case able to fail for the defect it was written for.

Found while triaging an unrelated red check on PR #3095: that diff changes attention and embedding gating in src/vt/ops.cpp and never calls the measured path, which is Threadpool::Run with an empty body.

主要言語
C++
スター
423
フォーク
53
平均マージ
1日 7時間
マージ済み PR(30日)
382

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

mudler/vllm.cpp のほかの issue

mudler/vllm.cpp の issue をすべて見る

似ている issue

C++ の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。