Hacktoberfest 2026:維護者為十月標記出來的 issue,仍然開放、適合新手。 瀏覽 Hacktoberfest issue

Extreme, load dependent slowness on multi-core CPUs with, for example, many iterations of a complex eigenproblem.

未關閉
#5,517 5 則留言 1 個 reaction 已指派 0 人 在 GitHub 檢視

維護者通常 1 天內回覆

還沒有人認領這個 Issue。

評估

難度
5/5
預估耗時
一週以上
新手友好度
25/100
Issue 類型
缺陷
描述清晰度
需要釐清
活躍度
停滯
技術堆疊
c, numpy
領域
performance

研究方向

首先,在報告的 OpenBLAS 與 OMP_NUM_THREADS 設定下,透過重複呼叫 numpy.linalg.eigh() 來重現附帶的測試案例。檢查偵錯器堆疊中顯示的 OpenBLAS exec_blas/gomp_barrier_wait_end() 路徑,然後比較存在背景 CPU 負載時的行為。完成的標準是找出並處理過度變慢的問題,同時不使小型特徵值問題的效能發生回歸。

由索引模型根據 Issue 內容生成。

描述

Very large numbers of calls to the symmetric complex eigenproblem via numpy.linalg.eigh() can have dramatic slowdowns due to multi-threading, especially when the number of cores is large and when there is significant cpu utilisation from other processes. The attached test case script runs around 130,000 iterations of 2x2 complex symmetric eigenproblem.

For example, on an AMD EPYC 7351P (16 cores/32 threads), RHEL10 OpenBLAS 0.3.28, with OMP_NUM_THREADS=1, the test completes in roughly 0.25s. With unlimited threads, the test completes in around 2.6s. With unlimited threads and a single thread background cpu hog ("cat /dev/zero >/dev/null"), the test completes in around 170s.

This incredible slowdown (over 500x in that first example!) also happens on other hardware/software configurations, but it seems to be less with fewer cores. On an Intel Core i7-1165G7 (4 cores/8 threads), Fedora 42 OpenBLAS 0.3.29, the same tests take roughly 0.125s, 0.4s, and 4-6s, respectively. By contrast more cores seem to make it worse; on an AMD Threadripper PRO 7965WX (24 cores/48 threads), RHEL9.6 OpenBLAS 0.3.26, the same tests take roughly 0.1s, 1.4s, and 800-900s respectively.

The very bad slowdowns seem to occur consistently when the combined number of OpenBLAS threads and background cpu-using threads exceeds the cpu thread count, although the situation is markedly worse when there is at least one CPU-hungry non-OpenBLAS background thread. For example, with no background jobs on the 32-thread EPYC 7351P with OMP_NUM_THREADS = 33, the test case takes around 27 seconds, whereas with two background cat jobs and OMP_NUM_THREADS = 31, the test case takes around 160-170s.

So why the horrendous slowdowns? First, numpy doesn't parallelize over the 130,000 different matrices. Instead, OpenBLAS seems to be trying to use the full thread complement of a large CPU to solve a 2x2 eigenproblem. One obvious fix would be to limit the number of threads involved to something on the order of the size of the matrix. However, the extreme slowdowns seen in this testcase are indicative of a deeper problem.

When I attach a debugger to the slow process, I consistently find it waiting in gomp_barrier_wait_end():

#0  0x00007f3688765cb6 in gomp_barrier_wait_end () from /lib64/libgomp.so.1
#1  0x00007f3688763fb1 in gomp_team_start () from /lib64/libgomp.so.1
#2  0x00007f368875a571 in GOMP_parallel () from /lib64/libgomp.so.1
#3  0x00007f3684be2681 in exec_blas () from /lib64/libopenblaso.so.0
#4  0x00007f3684a4d178 in zher2_thread_L () from /lib64/libopenblaso.so.0
#5  0x00007f36849c706d in zher2_ () from /lib64/libopenblaso.so.0
#6  0x00007f3686b33281 in zhetd2_ () from /lib64/libopenblaso.so.0
#7  0x00007f3686b351c9 in zhetrd_ () from /lib64/libopenblaso.so.0
#8  0x00007f3686b2c04e in zheevd_ () from /lib64/libopenblaso.so.0
#9  0x00007f350bf2f948 in void eigh_wrapper<npy_cdouble>(char, char, char**, long const*, long const*) ()
   from /usr/lib64/python3.9/site-packages/numpy/linalg/_umath_linalg.cpython-39-x86_64-linux-gnu.so

An OS expert I consulted with suggested based on the symptoms that the threads are likely "caravaning". Imagine a 32 lane (i.e. CPU core/thread) highway with 32 cars (OpenBLAS threads) and 1 tractor-trailer (background cat job). At the end of each eigenproblem, all 32 cars have to line up with each other (gomp_barrier_wait_end), but one of them is stuck behind the tractor-trailer and can't or won't go around, so it has to wait for the tractor-trailer to use up its scheduler quantum before the car can synchronize with the rest of the threads.

But then why can't the car just go around the tractor-trailer? The OpenBLAS threads should be otherwise idle, freeing their lanes for passing. CPU meters show 100% usage for all OpenBLAS threads, suggesting that yielding isn't happening or working (I know this is a tradeoff vs spamming the kernel; I tried setting the environment variable OPENBLAS_THREAD_TIMEOUT=1 but that didn't seem to have any effect). I don't think thread affinity is enabled in my default configuration.

In addition, the above explanation would imply that waiting out a scheduler quantum should be enough to recover the CPU. With modern Linux kernels running at HZ = 1000 and 130,000 eigenproblems, we would expect the delay to be upper bounded at 130s or a small multiple thereof. Yet the 48-thread CPU took as much as 900s.

In any case, a variety of issues probably derive from the same root cause of this slowdown. For example, #5383, perhaps #4496, #4651, #5469, etc. Hopefully this testcase is helpful. The OS expert I consulted thinks that it is very plausible that there might be OS scheduler issues contributing to these huge slowdowns.

主要語言
C
星號
7.6k
分支
1.7k
平均合併
1 天 13 小時
30 天內合併 PR
51

環境準備

這個專案沒有提供開發容器、Dockerfile 或貢獻指南,環境需要你自己搭建:先看它的 README,通用步驟見我們的新手貢獻指南。

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

OpenMathLib/OpenBLAS 的其他 Issue

查看 OpenMathLib/OpenBLAS 的全部 Issue

相似的 Issue

更多 C Issue

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。