metal-whisper: ggml_metal_library_init fails to compile embedded Metal shader library on macOS 26.5.2 (Apple M4 Pro) — flash-attn kernel's threadgroup half4x4 array rejected
Maintainer thường phản hồi trong vòng 3 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 38/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Ít trao đổi
- Công nghệ
- macos
- Lĩnh vực
- backend
Hướng nghiên cứu
Bắt đầu bằng cách tái hiện lỗi thông qua run.sh hoặc binary whisper, sau đó kiểm tra đường dẫn ggml_metal_library_init và khai báo shader flash-attention được nhúng có tên trong báo cáo. So sánh hành vi của compiler trên các đường dẫn macOS và phần cứng đã nêu, đồng thời xác minh rằng việc biên dịch library thất bại sẽ hiển thị thông tin chẩn đoán qua log của backend LocalAI thay vì chỉ trả về exitCode=2 và EOF.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Summary
metal-whisper fails to load any model on macOS 26.5.2 with Apple M4 Pro. ggml_metal_library_init fails to compile the backend's entire embedded Metal shader library, because one kernel (flash-attention) declares a threadgroup array of half4x4 matrices that the current Metal shader compiler on this OS/toolchain rejects. Since ggml compiles its whole shader library up front rather than per-kernel-on-demand, this one bad kernel makes Metal unusable for whisper on this machine — even for models/requests that never use flash attention.
This is the same failure class as the companion reports #11529 (https://github.com/mudler/LocalAI/issues/11529) and #11530 (https://github.com/mudler/LocalAI/issues/11530) (stablediffusion-ggml, also on macOS 26.5.2): a Metal pipeline/library compile failure that surfaces to the operator only as exitCode=2 and rpc error: code = Unavailable desc = error reading from server: EOF, with the actual Metal compiler diagnostic invisible unless you manually redirect the backend's stderr.
Environment
- LocalAI v4.8.2 (
5ff25d9d145e0a03a5b9a3559c620f1e1204ca6d) - Backend
metal-whisper, installed fromquay.io/go-skynet/local-ai-backends:latest-metal-darwin-arm64-whisper, digestsha256:4f18b4b228d2a750e74cf9f006e3789bb3d99da5b9d651e4fc1d8e55121346fa - Also reproduced on
metal-whisper-development(:master-metal-darwin-arm64-whisper), same failure - macOS 26.5.2 (build 25F84), Apple M4 Pro (Mac16,8)
- Model:
ggml-large-v3-turbo.bin(whisper-large-v3-turbo), validggmlmagic, 1.62 GB
What LocalAI reports
Aug 20 15:18:54 WARN Backend process exited unexpectedly id="whisper-large-turbo" address="127.0.0.1:50266" process="run.sh" exitCode="2"
Aug 20 15:18:54 ERROR Failed to load model modelID="whisper-large-turbo" error=failed to load model with internal loader: could not load model: rpc error: code = Unavailable desc = error reading from server: EOF backend="whisper"
No indication a Metal shader ever failed to compile.
What is actually happening
Capturing the backend's own stderr (by running run.sh/the whisper binary directly instead of through LocalAI) shows the real error:
[INFO ] whisper_init_with_params_no_state: use gpu = 1
[INFO ] whisper_init_with_params_no_state: flash attn = 1
[INFO ] ggml_metal_device_init: tensor API disabled for pre-M5 and pre-A19 devices
[INFO ] ggml_metal_library_init: using embedded metal library
[ERROR] ggml_metal_library_init: error: Error Domain=MTLLibraryErrorDomain Code=3
"program_source:14536:27: error: no matching constructor for initialization of
'threadgroup metal::half4x4[512]' (aka 'threadgroup matrix<half, 4, 4>[512]')
threadgroup half4x4 sk4x4[NK*DK16];
^
followed by ~450 lines of Metal compiler candidate-constructor diagnostics for metal::matrix, all stemming from the same threadgroup half4x4 sk4x4[NK*DK16] declaration (the flash-attention Metal kernel).
Because ggml/whisper.cpp compiles the entire embedded Metal shader source as one library (ggml_metal_library_init), this single kernel failing to compile takes down Metal initialization for the whole backend — even though flash_attention: "off" in the model's LocalAI YAML has no effect here (the library fails to compile before any per-request kernel selection happens).
Reproduction (bypassing LocalAI to isolate the backend)
cd ~/backends/metal-whisper
DYLD_LIBRARY_PATH="$(pwd)/lib" WHISPER_LIBRARY="$(pwd)/libgowhisper-fallback.so" \
./whisper -addr=127.0.0.1:59999 &
grpcurl -plaintext -proto backend.proto -d '{"Model":"ggml-large-v3-turbo.bin","ModelFile":"/path/to/ggml-large-v3-turbo.bin","Threads":14,"ContextSize":4096,"NBatch":512,"NGPULayers":99999999}' \
127.0.0.1:59999 backend.Backend/LoadModel
→ ERROR: Code: Unavailable, Message: error reading from server: EOF, and the backend's own stderr shows the Metal compile error above.
What would help
- Fix the ggml Metal shader source so the flash-attention kernel's
threadgroup half4x4 sk4x4[NK*DK16]array-of-matrices declaration compiles under current Xcode/macOS 26.x Metal shader compilers (this is presumably a ggml-upstream fix, shared withggml-org/whisper.cpp/ggml-org/llama.cpp, givensk4x4/NK*DK16look like shared flash-attention kernel naming). - Same asks as #11529 (https://github.com/mudler/LocalAI/issues/11529): don't let a failed Metal library/pipeline compile surface only as
exitCode=2+EOF— propagate the real error, and get the backend's stderr into the LocalAI log so operators don't have to bypass LocalAI entirely to find the actual cause. - Given
ggml_metal_device_init: tensor API disabled for pre-M5 and pre-A19 devicesis logged right before the failure, it's possible this shader path is only exercised/broken on non-M5 tensor-API-disabled devices — worth checking whether M5 Macs (where the tensor API is enabled) take a different code path and avoid this.
Related
- #11529 (https://github.com/mudler/LocalAI/issues/11529) (stablediffusion-ggml: Metal pipeline compile failure surfaces only as SIGSEGV/exitCode=2/EOF, same macOS 26.5.2)
- #11530 (https://github.com/mudler/LocalAI/issues/11530) (stablediffusion-ggml: missing Metal kernel for bf16, same root symptom pattern)
- Ngôn ngữ chính
- Go
- Star
- 49.2k
- Fork
- 4.5k
- Merge trung bình
- 1 ngày 7 giờ
- Pull request đã merge (30 ngày)
- 340
Chuẩn bị môi trường
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của mudler/LocalAI
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
mudler/LocalAI#11995 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 3 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
mudler/LocalAI#11991 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 3 ngày
-
fish-speech: make compile:true usable on Blackwell sm_121 by honouring the CUDA toolkit's ptxasĐang mởenhancement
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
mudler/LocalAI#11348 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 3 ngày
-
feat: add automatic MCP transport selection for 2024-11-05 / 2025-03-26 / 2025-06-18 vs 2025-11-25Đang mởenhancement
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 65/100
mudler/LocalAI#12262 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 3 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 25/100
Maintainer thường phản hồi trong vòng 3 ngày
Tất cả issue của mudler/LocalAI
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
rossoctl/context-guru#346 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 90/100
prime-radiant-inc/evener#2883 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
gravitational/teleport#69805 ·
Maintainer thường phản hồi trong vòng 11 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Under Poisson sampling, the `PLDAccountant` composes the inner event both before and after samplingĐang mở
Độ khó 2/5 Nửa ngày Mức phù hợp với người mới 78/100
google/differential-privacy#496 ·