Extreme, load dependent slowness on multi-core CPUs with, for example, many iterations of a complex eigenproblem.
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 25/100
- Tipo de issue
- Error
- Claridad
- Necesita aclaración
- Estado de actividad
- Estancado
- Stack tecnológico
- c, numpy
- Área
- performance
Línea de trabajo
Comienza reproduciendo el caso de prueba adjunto mediante llamadas repetidas a numpy.linalg.eigh() con las configuraciones de OpenBLAS y OMP_NUM_THREADS indicadas. Inspecciona la ruta de OpenBLAS exec_blas/gomp_barrier_wait_end() que se muestra en la pila del depurador y, después, compara el comportamiento con carga de CPU en segundo plano. Se considerará terminado cuando se haya identificado y solucionado la ralentización excesiva sin degradar el rendimiento en problemas de autovalores pequeños.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Very large numbers of calls to the symmetric complex eigenproblem via numpy.linalg.eigh() can have dramatic slowdowns due to multi-threading, especially when the number of cores is large and when there is significant cpu utilisation from other processes. The attached test case script runs around 130,000 iterations of 2x2 complex symmetric eigenproblem.
For example, on an AMD EPYC 7351P (16 cores/32 threads), RHEL10 OpenBLAS 0.3.28, with OMP_NUM_THREADS=1, the test completes in roughly 0.25s. With unlimited threads, the test completes in around 2.6s. With unlimited threads and a single thread background cpu hog ("cat /dev/zero >/dev/null"), the test completes in around 170s.
This incredible slowdown (over 500x in that first example!) also happens on other hardware/software configurations, but it seems to be less with fewer cores. On an Intel Core i7-1165G7 (4 cores/8 threads), Fedora 42 OpenBLAS 0.3.29, the same tests take roughly 0.125s, 0.4s, and 4-6s, respectively. By contrast more cores seem to make it worse; on an AMD Threadripper PRO 7965WX (24 cores/48 threads), RHEL9.6 OpenBLAS 0.3.26, the same tests take roughly 0.1s, 1.4s, and 800-900s respectively.
The very bad slowdowns seem to occur consistently when the combined number of OpenBLAS threads and background cpu-using threads exceeds the cpu thread count, although the situation is markedly worse when there is at least one CPU-hungry non-OpenBLAS background thread. For example, with no background jobs on the 32-thread EPYC 7351P with OMP_NUM_THREADS = 33, the test case takes around 27 seconds, whereas with two background cat jobs and OMP_NUM_THREADS = 31, the test case takes around 160-170s.
So why the horrendous slowdowns? First, numpy doesn't parallelize over the 130,000 different matrices. Instead, OpenBLAS seems to be trying to use the full thread complement of a large CPU to solve a 2x2 eigenproblem. One obvious fix would be to limit the number of threads involved to something on the order of the size of the matrix. However, the extreme slowdowns seen in this testcase are indicative of a deeper problem.
When I attach a debugger to the slow process, I consistently find it waiting in gomp_barrier_wait_end():
#0 0x00007f3688765cb6 in gomp_barrier_wait_end () from /lib64/libgomp.so.1
#1 0x00007f3688763fb1 in gomp_team_start () from /lib64/libgomp.so.1
#2 0x00007f368875a571 in GOMP_parallel () from /lib64/libgomp.so.1
#3 0x00007f3684be2681 in exec_blas () from /lib64/libopenblaso.so.0
#4 0x00007f3684a4d178 in zher2_thread_L () from /lib64/libopenblaso.so.0
#5 0x00007f36849c706d in zher2_ () from /lib64/libopenblaso.so.0
#6 0x00007f3686b33281 in zhetd2_ () from /lib64/libopenblaso.so.0
#7 0x00007f3686b351c9 in zhetrd_ () from /lib64/libopenblaso.so.0
#8 0x00007f3686b2c04e in zheevd_ () from /lib64/libopenblaso.so.0
#9 0x00007f350bf2f948 in void eigh_wrapper<npy_cdouble>(char, char, char**, long const*, long const*) ()
from /usr/lib64/python3.9/site-packages/numpy/linalg/_umath_linalg.cpython-39-x86_64-linux-gnu.so
An OS expert I consulted with suggested based on the symptoms that the threads are likely "caravaning". Imagine a 32 lane (i.e. CPU core/thread) highway with 32 cars (OpenBLAS threads) and 1 tractor-trailer (background cat job). At the end of each eigenproblem, all 32 cars have to line up with each other (gomp_barrier_wait_end), but one of them is stuck behind the tractor-trailer and can't or won't go around, so it has to wait for the tractor-trailer to use up its scheduler quantum before the car can synchronize with the rest of the threads.
But then why can't the car just go around the tractor-trailer? The OpenBLAS threads should be otherwise idle, freeing their lanes for passing. CPU meters show 100% usage for all OpenBLAS threads, suggesting that yielding isn't happening or working (I know this is a tradeoff vs spamming the kernel; I tried setting the environment variable OPENBLAS_THREAD_TIMEOUT=1 but that didn't seem to have any effect). I don't think thread affinity is enabled in my default configuration.
In addition, the above explanation would imply that waiting out a scheduler quantum should be enough to recover the CPU. With modern Linux kernels running at HZ = 1000 and 130,000 eigenproblems, we would expect the delay to be upper bounded at 130s or a small multiple thereof. Yet the 48-thread CPU took as much as 900s.
In any case, a variety of issues probably derive from the same root cause of this slowdown. For example, #5383, perhaps #4496, #4651, #5469, etc. Hopefully this testcase is helpful. The OS expert I consulted thinks that it is very plausible that there might be OS scheduler issues contributing to these huge slowdowns.
- Lenguaje dominante
- C
- Estrellas
- 7.6k
- Forks
- 1.7k
- Merge medio
- 1 d 6 h
- PR fusionados (30 d)
- 46
Preparar el entorno
Aún no hemos revisado los archivos de configuración de este proyecto. Empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de OpenMathLib/OpenBLAS
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 88/100
OpenMathLib/OpenBLAS#6062 · 2 comentarios ·
Los mantenedores suelen responder en 1 día
-
Missing cgroup awarenessAbierto
Dificultad 4/5 3-5 días Aptitud para principiantes 52/100
OpenMathLib/OpenBLAS#6059 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 4/5 3-5 días Aptitud para principiantes 48/100
OpenMathLib/OpenBLAS#6029 · 21 comentarios ·
Los mantenedores suelen responder en 1 día
-
Dificultad 3/5 1-2 días Aptitud para principiantes 68/100
OpenMathLib/OpenBLAS#6028 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 4/5 3-5 días Aptitud para principiantes 35/100
OpenMathLib/OpenBLAS#6005 · 21 comentarios · 2 reacciones ·
Los mantenedores suelen responder en 1 día
Todos los issues de OpenMathLib/OpenBLAS
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
ARM-software/sysarch-acs#556 · 1 comentario ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Los mantenedores suelen responder en 1 día
-
bug needs triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
netdata/netdata#24062 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100