Extreme, load dependent slowness on multi-core CPUs with, for example, many iterations of a complex eigenproblem.
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 25/100
- issue の種類
- バグ
- 明瞭さ
- 説明が足りない
- 活発さ
- 停滞
- 技術スタック
- c, numpy
- 領域
- performance
調査の方向性
報告されている OpenBLAS および OMP_NUM_THREADS の設定で、numpy.linalg.eigh() を繰り返し呼び出して、まず添付されたテストケースを再現します。次に、デバッガーのスタックに示されている OpenBLAS exec_blas/gomp_barrier_wait_end() のパスを調査し、バックグラウンドの CPU 負荷がある場合の挙動と比較します。小規模な固有値問題の性能を悪化させることなく、過度な低速化を特定して対処できれば完了です。
索引モデルが issue の本文から書いたものです。
説明
Very large numbers of calls to the symmetric complex eigenproblem via numpy.linalg.eigh() can have dramatic slowdowns due to multi-threading, especially when the number of cores is large and when there is significant cpu utilisation from other processes. The attached test case script runs around 130,000 iterations of 2x2 complex symmetric eigenproblem.
For example, on an AMD EPYC 7351P (16 cores/32 threads), RHEL10 OpenBLAS 0.3.28, with OMP_NUM_THREADS=1, the test completes in roughly 0.25s. With unlimited threads, the test completes in around 2.6s. With unlimited threads and a single thread background cpu hog ("cat /dev/zero >/dev/null"), the test completes in around 170s.
This incredible slowdown (over 500x in that first example!) also happens on other hardware/software configurations, but it seems to be less with fewer cores. On an Intel Core i7-1165G7 (4 cores/8 threads), Fedora 42 OpenBLAS 0.3.29, the same tests take roughly 0.125s, 0.4s, and 4-6s, respectively. By contrast more cores seem to make it worse; on an AMD Threadripper PRO 7965WX (24 cores/48 threads), RHEL9.6 OpenBLAS 0.3.26, the same tests take roughly 0.1s, 1.4s, and 800-900s respectively.
The very bad slowdowns seem to occur consistently when the combined number of OpenBLAS threads and background cpu-using threads exceeds the cpu thread count, although the situation is markedly worse when there is at least one CPU-hungry non-OpenBLAS background thread. For example, with no background jobs on the 32-thread EPYC 7351P with OMP_NUM_THREADS = 33, the test case takes around 27 seconds, whereas with two background cat jobs and OMP_NUM_THREADS = 31, the test case takes around 160-170s.
So why the horrendous slowdowns? First, numpy doesn't parallelize over the 130,000 different matrices. Instead, OpenBLAS seems to be trying to use the full thread complement of a large CPU to solve a 2x2 eigenproblem. One obvious fix would be to limit the number of threads involved to something on the order of the size of the matrix. However, the extreme slowdowns seen in this testcase are indicative of a deeper problem.
When I attach a debugger to the slow process, I consistently find it waiting in gomp_barrier_wait_end():
#0 0x00007f3688765cb6 in gomp_barrier_wait_end () from /lib64/libgomp.so.1
#1 0x00007f3688763fb1 in gomp_team_start () from /lib64/libgomp.so.1
#2 0x00007f368875a571 in GOMP_parallel () from /lib64/libgomp.so.1
#3 0x00007f3684be2681 in exec_blas () from /lib64/libopenblaso.so.0
#4 0x00007f3684a4d178 in zher2_thread_L () from /lib64/libopenblaso.so.0
#5 0x00007f36849c706d in zher2_ () from /lib64/libopenblaso.so.0
#6 0x00007f3686b33281 in zhetd2_ () from /lib64/libopenblaso.so.0
#7 0x00007f3686b351c9 in zhetrd_ () from /lib64/libopenblaso.so.0
#8 0x00007f3686b2c04e in zheevd_ () from /lib64/libopenblaso.so.0
#9 0x00007f350bf2f948 in void eigh_wrapper<npy_cdouble>(char, char, char**, long const*, long const*) ()
from /usr/lib64/python3.9/site-packages/numpy/linalg/_umath_linalg.cpython-39-x86_64-linux-gnu.so
An OS expert I consulted with suggested based on the symptoms that the threads are likely "caravaning". Imagine a 32 lane (i.e. CPU core/thread) highway with 32 cars (OpenBLAS threads) and 1 tractor-trailer (background cat job). At the end of each eigenproblem, all 32 cars have to line up with each other (gomp_barrier_wait_end), but one of them is stuck behind the tractor-trailer and can't or won't go around, so it has to wait for the tractor-trailer to use up its scheduler quantum before the car can synchronize with the rest of the threads.
But then why can't the car just go around the tractor-trailer? The OpenBLAS threads should be otherwise idle, freeing their lanes for passing. CPU meters show 100% usage for all OpenBLAS threads, suggesting that yielding isn't happening or working (I know this is a tradeoff vs spamming the kernel; I tried setting the environment variable OPENBLAS_THREAD_TIMEOUT=1 but that didn't seem to have any effect). I don't think thread affinity is enabled in my default configuration.
In addition, the above explanation would imply that waiting out a scheduler quantum should be enough to recover the CPU. With modern Linux kernels running at HZ = 1000 and 130,000 eigenproblems, we would expect the delay to be upper bounded at 130s or a small multiple thereof. Yet the 48-thread CPU took as much as 900s.
In any case, a variety of issues probably derive from the same root cause of this slowdown. For example, #5383, perhaps #4496, #4651, #5469, etc. Hopefully this testcase is helpful. The OS expert I consulted thinks that it is very plausible that there might be OS scheduler issues contributing to these huge slowdowns.
- 主要言語
- C
- スター
- 7.6k
- フォーク
- 1.7k
- 平均マージ
- 1日 6時間
- マージ済み PR(30日)
- 46
環境構築
このプロジェクトの環境構築ファイルはまだ確認していません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
OpenMathLib/OpenBLAS のほかの issue
-
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
OpenMathLib/OpenBLAS#6062 · コメント 2 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 4/5 3〜5日 初心者へのやさしさ 52/100
OpenMathLib/OpenBLAS#6059 ·
メンテナーはふだん 1 日以内に返信
-
難易度 4/5 3〜5日 初心者へのやさしさ 48/100
OpenMathLib/OpenBLAS#6029 · コメント 21 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 3/5 1〜2日 初心者へのやさしさ 68/100
OpenMathLib/OpenBLAS#6028 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 4/5 3〜5日 初心者へのやさしさ 35/100
OpenMathLib/OpenBLAS#6005 · コメント 21 件 · リアクション 2 件 ·
メンテナーはふだん 1 日以内に返信
OpenMathLib/OpenBLAS の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
ARM-software/sysarch-acs#556 · コメント 1 件 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
メンテナーはふだん 1 日以内に返信
-
bug needs triage
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
netdata/netdata#24062 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100