metal-whisper: ggml_metal_library_init fails to compile embedded Metal shader library on macOS 26.5.2 (Apple M4 Pro) — flash-attn kernel's threadgroup half4x4 array rejected
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 38/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Tranquilla
- Stack tecnologico
- macos
- Ambito
- backend
Direzione di ricerca
Inizia riproducendo il fallimento tramite run.sh o il binario whisper, quindi esamina il percorso ggml_metal_library_init e la dichiarazione dello shader flash-attention incorporato menzionata nel report. Confronta il comportamento del compilatore sui percorsi macOS e hardware indicati e verifica che una compilazione della libreria non riuscita esponga la relativa diagnosi nei log del backend LocalAI invece di restituire soltanto exitCode=2 e EOF.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Summary
metal-whisper fails to load any model on macOS 26.5.2 with Apple M4 Pro. ggml_metal_library_init fails to compile the backend's entire embedded Metal shader library, because one kernel (flash-attention) declares a threadgroup array of half4x4 matrices that the current Metal shader compiler on this OS/toolchain rejects. Since ggml compiles its whole shader library up front rather than per-kernel-on-demand, this one bad kernel makes Metal unusable for whisper on this machine — even for models/requests that never use flash attention.
This is the same failure class as the companion reports #11529 (https://github.com/mudler/LocalAI/issues/11529) and #11530 (https://github.com/mudler/LocalAI/issues/11530) (stablediffusion-ggml, also on macOS 26.5.2): a Metal pipeline/library compile failure that surfaces to the operator only as exitCode=2 and rpc error: code = Unavailable desc = error reading from server: EOF, with the actual Metal compiler diagnostic invisible unless you manually redirect the backend's stderr.
Environment
- LocalAI v4.8.2 (
5ff25d9d145e0a03a5b9a3559c620f1e1204ca6d) - Backend
metal-whisper, installed fromquay.io/go-skynet/local-ai-backends:latest-metal-darwin-arm64-whisper, digestsha256:4f18b4b228d2a750e74cf9f006e3789bb3d99da5b9d651e4fc1d8e55121346fa - Also reproduced on
metal-whisper-development(:master-metal-darwin-arm64-whisper), same failure - macOS 26.5.2 (build 25F84), Apple M4 Pro (Mac16,8)
- Model:
ggml-large-v3-turbo.bin(whisper-large-v3-turbo), validggmlmagic, 1.62 GB
What LocalAI reports
Aug 20 15:18:54 WARN Backend process exited unexpectedly id="whisper-large-turbo" address="127.0.0.1:50266" process="run.sh" exitCode="2"
Aug 20 15:18:54 ERROR Failed to load model modelID="whisper-large-turbo" error=failed to load model with internal loader: could not load model: rpc error: code = Unavailable desc = error reading from server: EOF backend="whisper"
No indication a Metal shader ever failed to compile.
What is actually happening
Capturing the backend's own stderr (by running run.sh/the whisper binary directly instead of through LocalAI) shows the real error:
[INFO ] whisper_init_with_params_no_state: use gpu = 1
[INFO ] whisper_init_with_params_no_state: flash attn = 1
[INFO ] ggml_metal_device_init: tensor API disabled for pre-M5 and pre-A19 devices
[INFO ] ggml_metal_library_init: using embedded metal library
[ERROR] ggml_metal_library_init: error: Error Domain=MTLLibraryErrorDomain Code=3
"program_source:14536:27: error: no matching constructor for initialization of
'threadgroup metal::half4x4[512]' (aka 'threadgroup matrix<half, 4, 4>[512]')
threadgroup half4x4 sk4x4[NK*DK16];
^
followed by ~450 lines of Metal compiler candidate-constructor diagnostics for metal::matrix, all stemming from the same threadgroup half4x4 sk4x4[NK*DK16] declaration (the flash-attention Metal kernel).
Because ggml/whisper.cpp compiles the entire embedded Metal shader source as one library (ggml_metal_library_init), this single kernel failing to compile takes down Metal initialization for the whole backend — even though flash_attention: "off" in the model's LocalAI YAML has no effect here (the library fails to compile before any per-request kernel selection happens).
Reproduction (bypassing LocalAI to isolate the backend)
cd ~/backends/metal-whisper
DYLD_LIBRARY_PATH="$(pwd)/lib" WHISPER_LIBRARY="$(pwd)/libgowhisper-fallback.so" \
./whisper -addr=127.0.0.1:59999 &
grpcurl -plaintext -proto backend.proto -d '{"Model":"ggml-large-v3-turbo.bin","ModelFile":"/path/to/ggml-large-v3-turbo.bin","Threads":14,"ContextSize":4096,"NBatch":512,"NGPULayers":99999999}' \
127.0.0.1:59999 backend.Backend/LoadModel
→ ERROR: Code: Unavailable, Message: error reading from server: EOF, and the backend's own stderr shows the Metal compile error above.
What would help
- Fix the ggml Metal shader source so the flash-attention kernel's
threadgroup half4x4 sk4x4[NK*DK16]array-of-matrices declaration compiles under current Xcode/macOS 26.x Metal shader compilers (this is presumably a ggml-upstream fix, shared withggml-org/whisper.cpp/ggml-org/llama.cpp, givensk4x4/NK*DK16look like shared flash-attention kernel naming). - Same asks as #11529 (https://github.com/mudler/LocalAI/issues/11529): don't let a failed Metal library/pipeline compile surface only as
exitCode=2+EOF— propagate the real error, and get the backend's stderr into the LocalAI log so operators don't have to bypass LocalAI entirely to find the actual cause. - Given
ggml_metal_device_init: tensor API disabled for pre-M5 and pre-A19 devicesis logged right before the failure, it's possible this shader path is only exercised/broken on non-M5 tensor-API-disabled devices — worth checking whether M5 Macs (where the tensor API is enabled) take a different code path and avoid this.
Related
- #11529 (https://github.com/mudler/LocalAI/issues/11529) (stablediffusion-ggml: Metal pipeline compile failure surfaces only as SIGSEGV/exitCode=2/EOF, same macOS 26.5.2)
- #11530 (https://github.com/mudler/LocalAI/issues/11530) (stablediffusion-ggml: missing Metal kernel for bf16, same root symptom pattern)
- Lingua principale
- Go
- Stelle
- 49.2k
- Fork
- 4.5k
- Merge medio
- 1g 7h
- PR unite (30g)
- 357
Preparare l'ambiente
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di mudler/LocalAI
-
bug unconfirmed
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
mudler/LocalAI#12337 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
mudler/LocalAI#11995 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
mudler/LocalAI#11991 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
fish-speech: make compile:true usable on Blackwell sm_121 by honouring the CUDA toolkit's ptxasApertaenhancement
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
mudler/LocalAI#11348 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
bug unconfirmed
Difficoltà 4/5 3-5 giorni Idoneità per principianti 70/100
mudler/LocalAI#12331 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di mudler/LocalAI
Issue simili
-
agentic-workflows
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
I maintainer di solito rispondono entro 1 giorno
-
priority/4/normal status/needs-triage type/bug/unconfirmed
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
authelia/authelia#13292 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
blinklabs-io/actions#138 ·
I maintainer di solito rispondono entro 1 giorno
-
[UI] AlbumDetails collapses multi-genre list to single primary genre on viewports < lg breakpointAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
I maintainer di solito rispondono entro 1 giorno