Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

arm64 SME2 (Apple M4): SGEMM/DGEMM well below the hardware, level-3 routines on NEON, M4 Pro not detected

Abierto
#6,073 2 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
35/100
Tipo de issue
Error
Claridad
Bastante claro
Estado de actividad
Activo
Stack tecnológico
c

Línea de trabajo

Start with the SME [s/d]GEMM kernels, the level-3 driver, and getarch, using the references to issues #5971 and #5011 for existing context. Reproduce the reported M4 Pro benchmarks and detection behavior; done means M4 Pro is recognized and the affected level-3 routines use appropriate SME paths with performance approaching the stated hardware and Accelerate comparisons.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

On Apple M4 (SME2, 512-bit streaming vector length), develop (63d7f22) leaves most of the SME unit unused:

  • SGEMM/DGEMM: sme_[sd]gemm_kernel (#5971) runs at about 0.6x of Accelerate on one thread, and on one
    thread whatever the thread count, while the M4 Pro has two SME units (one per performance cluster).
    fp32 geometric means over the 24 workloads of Deng et al. (arXiv:2512.21473), row-major / column-major, GFLOPS:

    Accelerate develop
    SGEMM, 1 thread 1127 / 1185 707 / 766
    SGEMM, default threads 2364 / 2445 702 / 766
    DGEMM, 1 thread 321 / 344 275 / 283
  • SYMM, SYRK, SYR2K, TRMM, TRSM run through the level-3 driver on NEON kernels: about 100-115 GFLOPS in fp32
    and 53-58 in fp64 at n = 1024 on one thread, against 791-1758 and 265-433 for Accelerate. The SME kernel cannot
    simply be plugged into the driver, because SYMM/TRMM share the GEMM unroll sizes (where #5011 stopped).

  • Detection: the M4 Pro (hw.cpufamily 0x17d5b93a) is not recognised by getarch, so a build without
    TARGET falls back to ARMV8 without any SME code.

Measured on an M4 Pro, macOS 27, Apple clang 21; harness and raw data in https://github.com/tesch1/mtgemm-a.

Prepared with AI coding assistance.

Lenguaje dominante
C
Estrellas
7.6k
Forks
1.7k
Merge medio
1 d 5 h
PR fusionados (30 d)
56

Preparar el entorno

Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de OpenMathLib/OpenBLAS

Todos los issues de OpenMathLib/OpenBLAS

Issues similares

Más issues de C

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.