Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

metal-whisper: ggml_metal_library_init fails to compile embedded Metal shader library on macOS 26.5.2 (Apple M4 Pro) — flash-attn kernel's threadgroup half4x4 array rejected

未关闭
#11,624 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
5/5
预计耗时
一周以上
新手友好度
38/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
冷清
技术栈
macos
领域
backend

调研方向

首先通过 run.sh 或 whisper 二进制文件重现该失败,然后检查 ggml_metal_library_init 路径以及报告中提到的嵌入式 flash-attention shader 声明。比较指定的 macOS 和硬件路径上的编译器行为,并验证库编译失败时能够通过 LocalAI 后端日志公开其诊断信息,而不是仅显示 exitCode=2 和 EOF。

由索引模型根据 Issue 内容生成。

描述

area/backends bug confirmed os/macOS upstream issue

Summary

metal-whisper fails to load any model on macOS 26.5.2 with Apple M4 Pro. ggml_metal_library_init fails to compile the backend's entire embedded Metal shader library, because one kernel (flash-attention) declares a threadgroup array of half4x4 matrices that the current Metal shader compiler on this OS/toolchain rejects. Since ggml compiles its whole shader library up front rather than per-kernel-on-demand, this one bad kernel makes Metal unusable for whisper on this machine — even for models/requests that never use flash attention.

This is the same failure class as the companion reports #11529 (https://github.com/mudler/LocalAI/issues/11529) and #11530 (https://github.com/mudler/LocalAI/issues/11530) (stablediffusion-ggml, also on macOS 26.5.2): a Metal pipeline/library compile failure that surfaces to the operator only as exitCode=2 and rpc error: code = Unavailable desc = error reading from server: EOF, with the actual Metal compiler diagnostic invisible unless you manually redirect the backend's stderr.

Environment

  • LocalAI v4.8.2 (5ff25d9d145e0a03a5b9a3559c620f1e1204ca6d)
  • Backend metal-whisper, installed from quay.io/go-skynet/local-ai-backends:latest-metal-darwin-arm64-whisper, digest sha256:4f18b4b228d2a750e74cf9f006e3789bb3d99da5b9d651e4fc1d8e55121346fa
  • Also reproduced on metal-whisper-development (:master-metal-darwin-arm64-whisper), same failure
  • macOS 26.5.2 (build 25F84), Apple M4 Pro (Mac16,8)
  • Model: ggml-large-v3-turbo.bin (whisper-large-v3-turbo), valid ggml magic, 1.62 GB

What LocalAI reports

Aug 20 15:18:54 WARN  Backend process exited unexpectedly id="whisper-large-turbo" address="127.0.0.1:50266" process="run.sh" exitCode="2"
Aug 20 15:18:54 ERROR Failed to load model modelID="whisper-large-turbo" error=failed to load model with internal loader: could not load model: rpc error: code = Unavailable desc = error reading from server: EOF backend="whisper"

No indication a Metal shader ever failed to compile.

What is actually happening

Capturing the backend's own stderr (by running run.sh/the whisper binary directly instead of through LocalAI) shows the real error:

[INFO ] whisper_init_with_params_no_state: use gpu    = 1
[INFO ] whisper_init_with_params_no_state: flash attn = 1
[INFO ] ggml_metal_device_init: tensor API disabled for pre-M5 and pre-A19 devices
[INFO ] ggml_metal_library_init: using embedded metal library
[ERROR] ggml_metal_library_init: error: Error Domain=MTLLibraryErrorDomain Code=3
"program_source:14536:27: error: no matching constructor for initialization of
'threadgroup metal::half4x4[512]' (aka 'threadgroup matrix<half, 4, 4>[512]')
    threadgroup half4x4   sk4x4[NK*DK16];
                          ^

followed by ~450 lines of Metal compiler candidate-constructor diagnostics for metal::matrix, all stemming from the same threadgroup half4x4 sk4x4[NK*DK16] declaration (the flash-attention Metal kernel).

Because ggml/whisper.cpp compiles the entire embedded Metal shader source as one library (ggml_metal_library_init), this single kernel failing to compile takes down Metal initialization for the whole backend — even though flash_attention: "off" in the model's LocalAI YAML has no effect here (the library fails to compile before any per-request kernel selection happens).

Reproduction (bypassing LocalAI to isolate the backend)

cd ~/backends/metal-whisper
DYLD_LIBRARY_PATH="$(pwd)/lib" WHISPER_LIBRARY="$(pwd)/libgowhisper-fallback.so" \
  ./whisper -addr=127.0.0.1:59999 &
grpcurl -plaintext -proto backend.proto -d '{"Model":"ggml-large-v3-turbo.bin","ModelFile":"/path/to/ggml-large-v3-turbo.bin","Threads":14,"ContextSize":4096,"NBatch":512,"NGPULayers":99999999}' \
  127.0.0.1:59999 backend.Backend/LoadModel

→ ERROR: Code: Unavailable, Message: error reading from server: EOF, and the backend's own stderr shows the Metal compile error above.

What would help

  1. Fix the ggml Metal shader source so the flash-attention kernel's threadgroup half4x4 sk4x4[NK*DK16] array-of-matrices declaration compiles under current Xcode/macOS 26.x Metal shader compilers (this is presumably a ggml-upstream fix, shared with ggml-org/whisper.cpp / ggml-org/llama.cpp, given sk4x4/NK*DK16 look like shared flash-attention kernel naming).
  2. Same asks as #11529 (https://github.com/mudler/LocalAI/issues/11529): don't let a failed Metal library/pipeline compile surface only as exitCode=2 + EOF — propagate the real error, and get the backend's stderr into the LocalAI log so operators don't have to bypass LocalAI entirely to find the actual cause.
  3. Given ggml_metal_device_init: tensor API disabled for pre-M5 and pre-A19 devices is logged right before the failure, it's possible this shader path is only exercised/broken on non-M5 tensor-API-disabled devices — worth checking whether M5 Macs (where the tensor API is enabled) take a different code path and avoid this.

Related

主要语言
Go
星标
49.2k
派生
4.5k
平均合并
1 天 7 小时
30 天内合并 PR
362

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

mudler/LocalAI 的其他 Issue

查看 mudler/LocalAI 的全部 Issue

相似的 Issue

更多 Go Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。