Hacktoberfest 2026 : les issues que les mainteneurs ont marquées pour octobre, ouvertes et accessibles aux débutants. Parcourir les issues Hacktoberfest

arm64 SME2 (Apple M4): SGEMM/DGEMM well below the hardware, level-3 routines on NEON, M4 Pro not detected

Ouverte
#6,073 1 commentaire 0 réactions 0 personnes assignées Voir sur GitHub

Les mainteneurs répondent en général sous 1 jour

Personne n'a encore pris cette issue.

Évaluation

Difficulté
5/5
Temps estimé
Plus d'une semaine
Accessibilité débutants
35/100
Type d'issue
Bug
Clarté
Plutôt claire
Activité
Active
Stack technique
c
Domaine
backend, performance

Piste de recherche

Start with the SME [s/d]GEMM kernels, the level-3 driver, and getarch, using the references to issues #5971 and #5011 for existing context. Reproduce the reported M4 Pro benchmarks and detection behavior; done means M4 Pro is recognized and the affected level-3 routines use appropriate SME paths with performance approaching the stated hardware and Accelerate comparisons.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Description

On Apple M4 (SME2, 512-bit streaming vector length), develop (63d7f22) leaves most of the SME unit unused:

  • SGEMM/DGEMM: sme_[sd]gemm_kernel (#5971) runs at about 0.6x of Accelerate on one thread, and on one
    thread whatever the thread count, while the M4 Pro has two SME units (one per performance cluster).
    fp32 geometric means over the 24 workloads of Deng et al. (arXiv:2512.21473), row-major / column-major, GFLOPS:

    Accelerate develop
    SGEMM, 1 thread 1127 / 1185 707 / 766
    SGEMM, default threads 2364 / 2445 702 / 766
    DGEMM, 1 thread 321 / 344 275 / 283
  • SYMM, SYRK, SYR2K, TRMM, TRSM run through the level-3 driver on NEON kernels: about 100-115 GFLOPS in fp32
    and 53-58 in fp64 at n = 1024 on one thread, against 791-1758 and 265-433 for Accelerate. The SME kernel cannot
    simply be plugged into the driver, because SYMM/TRMM share the GEMM unroll sizes (where #5011 stopped).

  • Detection: the M4 Pro (hw.cpufamily 0x17d5b93a) is not recognised by getarch, so a build without
    TARGET falls back to ARMV8 without any SME code.

Measured on an M4 Pro, macOS 27, Apple clang 21; harness and raw data in https://github.com/tesch1/mtgemm-a.

Prepared with AI coding assistance.

Langage dominant
C
Étoiles
7.6k
Forks
1.7k
Merge moyen
1 j 8 h
PR mergées (30 j)
48

Préparer son environnement

Ce projet ne fournit ni conteneur de développement, ni Dockerfile, ni guide de contribution : l'installation est à votre charge. Commencez par son README, et consultez notre guide de la première contribution pour les étapes générales.

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Autres issues de OpenMathLib/OpenBLAS

Toutes les issues de OpenMathLib/OpenBLAS

Issues similaires

Plus d'issues C

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.