Extreme, load dependent slowness on multi-core CPUs with, for example, many iterations of a complex eigenproblem.
Les mainteneurs répondent en général sous 1 jour
Personne n'a encore pris cette issue.
Évaluation
- Difficulté
- 5/5
- Temps estimé
- Plus d'une semaine
- Accessibilité débutants
- 25/100
- Type d'issue
- Bug
- Clarté
- À clarifier
- Activité
- À l'abandon
- Stack technique
- c, numpy
- Domaine
- performance
Piste de recherche
Commencez par reproduire le cas de test joint à l’aide d’appels répétés à numpy.linalg.eigh() avec les paramètres OpenBLAS et OMP_NUM_THREADS indiqués. Examinez le chemin OpenBLAS exec_blas/gomp_barrier_wait_end() affiché dans la pile du débogueur, puis comparez le comportement avec une charge CPU en arrière-plan. Le travail sera considéré comme terminé lorsque le ralentissement excessif aura été identifié et corrigé sans dégrader les performances pour les petits problèmes de valeurs propres.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Description
Very large numbers of calls to the symmetric complex eigenproblem via numpy.linalg.eigh() can have dramatic slowdowns due to multi-threading, especially when the number of cores is large and when there is significant cpu utilisation from other processes. The attached test case script runs around 130,000 iterations of 2x2 complex symmetric eigenproblem.
For example, on an AMD EPYC 7351P (16 cores/32 threads), RHEL10 OpenBLAS 0.3.28, with OMP_NUM_THREADS=1, the test completes in roughly 0.25s. With unlimited threads, the test completes in around 2.6s. With unlimited threads and a single thread background cpu hog ("cat /dev/zero >/dev/null"), the test completes in around 170s.
This incredible slowdown (over 500x in that first example!) also happens on other hardware/software configurations, but it seems to be less with fewer cores. On an Intel Core i7-1165G7 (4 cores/8 threads), Fedora 42 OpenBLAS 0.3.29, the same tests take roughly 0.125s, 0.4s, and 4-6s, respectively. By contrast more cores seem to make it worse; on an AMD Threadripper PRO 7965WX (24 cores/48 threads), RHEL9.6 OpenBLAS 0.3.26, the same tests take roughly 0.1s, 1.4s, and 800-900s respectively.
The very bad slowdowns seem to occur consistently when the combined number of OpenBLAS threads and background cpu-using threads exceeds the cpu thread count, although the situation is markedly worse when there is at least one CPU-hungry non-OpenBLAS background thread. For example, with no background jobs on the 32-thread EPYC 7351P with OMP_NUM_THREADS = 33, the test case takes around 27 seconds, whereas with two background cat jobs and OMP_NUM_THREADS = 31, the test case takes around 160-170s.
So why the horrendous slowdowns? First, numpy doesn't parallelize over the 130,000 different matrices. Instead, OpenBLAS seems to be trying to use the full thread complement of a large CPU to solve a 2x2 eigenproblem. One obvious fix would be to limit the number of threads involved to something on the order of the size of the matrix. However, the extreme slowdowns seen in this testcase are indicative of a deeper problem.
When I attach a debugger to the slow process, I consistently find it waiting in gomp_barrier_wait_end():
#0 0x00007f3688765cb6 in gomp_barrier_wait_end () from /lib64/libgomp.so.1
#1 0x00007f3688763fb1 in gomp_team_start () from /lib64/libgomp.so.1
#2 0x00007f368875a571 in GOMP_parallel () from /lib64/libgomp.so.1
#3 0x00007f3684be2681 in exec_blas () from /lib64/libopenblaso.so.0
#4 0x00007f3684a4d178 in zher2_thread_L () from /lib64/libopenblaso.so.0
#5 0x00007f36849c706d in zher2_ () from /lib64/libopenblaso.so.0
#6 0x00007f3686b33281 in zhetd2_ () from /lib64/libopenblaso.so.0
#7 0x00007f3686b351c9 in zhetrd_ () from /lib64/libopenblaso.so.0
#8 0x00007f3686b2c04e in zheevd_ () from /lib64/libopenblaso.so.0
#9 0x00007f350bf2f948 in void eigh_wrapper<npy_cdouble>(char, char, char**, long const*, long const*) ()
from /usr/lib64/python3.9/site-packages/numpy/linalg/_umath_linalg.cpython-39-x86_64-linux-gnu.so
An OS expert I consulted with suggested based on the symptoms that the threads are likely "caravaning". Imagine a 32 lane (i.e. CPU core/thread) highway with 32 cars (OpenBLAS threads) and 1 tractor-trailer (background cat job). At the end of each eigenproblem, all 32 cars have to line up with each other (gomp_barrier_wait_end), but one of them is stuck behind the tractor-trailer and can't or won't go around, so it has to wait for the tractor-trailer to use up its scheduler quantum before the car can synchronize with the rest of the threads.
But then why can't the car just go around the tractor-trailer? The OpenBLAS threads should be otherwise idle, freeing their lanes for passing. CPU meters show 100% usage for all OpenBLAS threads, suggesting that yielding isn't happening or working (I know this is a tradeoff vs spamming the kernel; I tried setting the environment variable OPENBLAS_THREAD_TIMEOUT=1 but that didn't seem to have any effect). I don't think thread affinity is enabled in my default configuration.
In addition, the above explanation would imply that waiting out a scheduler quantum should be enough to recover the CPU. With modern Linux kernels running at HZ = 1000 and 130,000 eigenproblems, we would expect the delay to be upper bounded at 130s or a small multiple thereof. Yet the 48-thread CPU took as much as 900s.
In any case, a variety of issues probably derive from the same root cause of this slowdown. For example, #5383, perhaps #4496, #4651, #5469, etc. Hopefully this testcase is helpful. The OS expert I consulted thinks that it is very plausible that there might be OS scheduler issues contributing to these huge slowdowns.
- Langage dominant
- C
- Étoiles
- 7.6k
- Forks
- 1.7k
- Merge moyen
- 1 j 6 h
- PR mergées (30 j)
- 46
Préparer son environnement
Ce projet ne fournit ni conteneur de développement, ni Dockerfile, ni guide de contribution : l'installation est à votre charge. Commencez par son README, et consultez notre guide de la première contribution pour les étapes générales.
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Autres issues de OpenMathLib/OpenBLAS
-
Difficulté 3/5 1-2 jours Accessibilité débutants 68/100
OpenMathLib/OpenBLAS#6069 ·
Les mainteneurs répondent en général sous 1 jour
-
Missing cgroup awarenessOuverte
Difficulté 4/5 3-5 jours Accessibilité débutants 52/100
OpenMathLib/OpenBLAS#6059 ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 4/5 3-5 jours Accessibilité débutants 48/100
OpenMathLib/OpenBLAS#6029 · 21 commentaires ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 3/5 1-2 jours Accessibilité débutants 68/100
OpenMathLib/OpenBLAS#6028 · 1 commentaire ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 4/5 3-5 jours Accessibilité débutants 35/100
OpenMathLib/OpenBLAS#6005 · 21 commentaires · 2 réactions ·
Les mainteneurs répondent en général sous 1 jour
Toutes les issues de OpenMathLib/OpenBLAS
Issues similaires
-
Difficulté 2/5 1-3 heures Accessibilité débutants 90/100
BasedHardware/omi#19711 ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 2/5 1-3 heures Accessibilité débutants 85/100
microsoft/ebpf-for-windows#5604 ·
Les mainteneurs répondent en général sous 3 jours
-
Difficulté 2/5 1-3 heures Accessibilité débutants 68/100
trezor/trezor-firmware#7985 ·
Les mainteneurs répondent en général sous 2 jours
-
Difficulté 2/5 1-3 heures Accessibilité débutants 84/100
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 2/5 1-3 heures Accessibilité débutants 70/100
Les mainteneurs répondent en général sous 2 jours