Hacktoberfest 2026: as issues que os mantenedores marcaram para outubro, abertas e boas para iniciantes. Ver issues do Hacktoberfest

arm64 SME2 (Apple M4): SGEMM/DGEMM well below the hardware, level-3 routines on NEON, M4 Pro not detected

Aberta
#6,073 2 comentários 0 reações 0 responsáveis Ver no GitHub

Mantenedores costumam responder em até 1 dia

Ninguém assumiu esta issue ainda.

Avaliação

Dificuldade
5/5
Tempo estimado
Mais de uma semana
Facilidade para iniciantes
35/100
Tipo de issue
Bug
Clareza
Razoavelmente clara
Status de atividade
Ativa
Stack de tecnologia
c
Domínio
backend, performance

Direção de pesquisa

Start with the SME [s/d]GEMM kernels, the level-3 driver, and getarch, using the references to issues #5971 and #5011 for existing context. Reproduce the reported M4 Pro benchmarks and detection behavior; done means M4 Pro is recognized and the affected level-3 routines use appropriate SME paths with performance approaching the stated hardware and Accelerate comparisons.

Escrita pelo modelo de indexação a partir do texto da issue.

Descrição

On Apple M4 (SME2, 512-bit streaming vector length), develop (63d7f22) leaves most of the SME unit unused:

  • SGEMM/DGEMM: sme_[sd]gemm_kernel (#5971) runs at about 0.6x of Accelerate on one thread, and on one
    thread whatever the thread count, while the M4 Pro has two SME units (one per performance cluster).
    fp32 geometric means over the 24 workloads of Deng et al. (arXiv:2512.21473), row-major / column-major, GFLOPS:

    Accelerate develop
    SGEMM, 1 thread 1127 / 1185 707 / 766
    SGEMM, default threads 2364 / 2445 702 / 766
    DGEMM, 1 thread 321 / 344 275 / 283
  • SYMM, SYRK, SYR2K, TRMM, TRSM run through the level-3 driver on NEON kernels: about 100-115 GFLOPS in fp32
    and 53-58 in fp64 at n = 1024 on one thread, against 791-1758 and 265-433 for Accelerate. The SME kernel cannot
    simply be plugged into the driver, because SYMM/TRMM share the GEMM unroll sizes (where #5011 stopped).

  • Detection: the M4 Pro (hw.cpufamily 0x17d5b93a) is not recognised by getarch, so a build without
    TARGET falls back to ARMV8 without any SME code.

Measured on an M4 Pro, macOS 27, Apple clang 21; harness and raw data in https://github.com/tesch1/mtgemm-a.

Prepared with AI coding assistance.

Linguagem predominante
C
Estrelas
7.6k
Forks
1.7k
Merge médio
1d 7h
PRs com merge (30d)
54

Preparar o ambiente

Este projeto não oferece contêiner de desenvolvimento, Dockerfile nem guia de contribuição, então a configuração fica por sua conta: comece pelo README e veja nosso guia da primeira contribuição para os passos gerais.

Primeiros passos

  1. Leia a issue inteira e depois o guia de contribuição do projeto.
  2. Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
  3. Faça um fork do repositório e trabalhe em uma branch.
  4. Abra um pull request que referencie o número da issue.

Mais de OpenMathLib/OpenBLAS

Todas as issues de OpenMathLib/OpenBLAS

Issues semelhantes

Mais issues de C

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.