cuda13-llama-cpp backend crashes loading gemma4 GGUF: "wrong number of tensors; expected 2131, got 720"
メンテナーはふだん 3 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 45/100
- issue の種類
- バグ
- 明瞭さ
- おおむね明確
- 活発さ
- 静か
- 領域
- backend
調査の方向性
cuda13-llama-cpp のバックエンド読み込みパスから始め、提供された gemma4 GGUF と POST /v1/chat/completions リクエストを使って失敗を再現します。loader がテンソル数の不一致とバックエンドの終了をどのように処理するか確認します。読み込みに成功するか、プロセスが終了せずにサポートされていないアーキテクチャを示すクリーンなエラーになることが完了条件です。
索引モデルが issue の本文から書いたものです。
説明
Description
Loading a gemma4 (Google Gemma 4, e.g. gemma4:e4b) architecture GGUF via the
cuda13-llama-cpp backend fails to load, with the backend process crashing
(exitCode -1) after emitting a tensor-count mismatch from llama.cpp's loader.
This blocks migrating any Ollama-served gemma4-family model onto LocalAI's
llama-cpp backend, since llama.cpp itself is where the incompatibility lives (see
"Root cause" below) — LocalAI is just the messenger here, but the crash-on-load
behavior (full backend process death instead of a clean load-time model
validation error) is a LocalAI/backend robustness gap worth tracking regardless
of when upstream llama.cpp lands full gemma4 support.
LocalAI version
v4.8.2 (5ff25d9d145e0a03a5b9a3559c620f1e1204ca6d), running via Docker image
localai/localai:latest-gpu-nvidia-cuda-13.
Backend: cuda13-llama-cpp, installed today via the gallery from
quay.io/go-skynet/local-ai-backends:latest-gpu-nvidia-cuda-13-llama-cpp
(digest sha256:389b92f825136696f85c548b9dcda65c87a224261d15f0089772dd1d40d5fda0).
Environment
- Docker container
localai/localai:latest-gpu-nvidia-cuda-13,--gpus all,
NVIDIA RTX 5080,nvidia-container-toolkitGPU passthrough verified working
(other models — a llama-3.2-1b test model and a Qwen3.6-35B-A3B MoE model —
load and infer correctly on this same instance via the samecuda13-llama-cpp
backend). - Host: Linux x86_64.
Model / GGUF
general.architecture(read directly from GGUF header):gemma4- GGUF header:
tensor_count = 2131,kv_count = 55, GGUF version 3 - This is the exact GGUF blob Ollama 0.32.6 ships/uses for its
gemma4:e4btag
(extracted byte-identical, sha256-verified, directly from Ollama's own blob
store — same file Ollama's own bundled llama-server process currently loads
and serves successfully with--no-jinja --chat-template chatml). - Quant: Q4_K_M, ~9.6GB file size, ~8.19B params.
Steps to reproduce
- Place the GGUF under LocalAI's models directory.
- Model YAML:
name: gemma4-e4b
backend: cuda13-llama-cpp
f16: true
mmap: true
context_size: 4096
parameters:
model: gemma4-e4b-q4_k_m.gguf
options:
- jinja
POST /v1/chat/completionswith"model": "gemma4-e4b".
Expected behavior
Either the model loads and serves a completion, or LocalAI returns a clean
"unsupported architecture" error without killing the backend process.
Actual behavior
Backend process crashes immediately on load attempt (exitCode -1), and the API
returns a 500 with this llama.cpp loader error:
Failed to load model: /models/gemma4-e4b-q4_k_m.gguf. Error: llama_model_load:
error loading model: done_getting_tensors: wrong number of tensors; expected
2131, got 720; llama_model_load_from_file_impl: failed to load model;
llama_model_load: error loading model: done_getting_tensors: wrong number of
tensors; expected 2131, got 720; llama_model_load_from_file_impl: failed to
load model
Full log excerpt:
INFO BackendLoader starting modelID=gemma4-e4b backend=cuda13-llama-cpp model=gemma4-e4b-q4_k_m.gguf
INFO effective runtime tuning ... context=4096 n_batch=2048 n_gpu_layers=99999999 parallel=1 flash_attention=auto f16=true
WARN Backend process exited unexpectedly id=gemma4-e4b address=127.0.0.1:34875 process=run.sh exitCode=-1
ERROR Failed to load model modelID=gemma4-e4b error=failed to load model with internal loader: could not load model: rpc error: code = Internal desc = Failed to load model: ... (as above)
Root cause (as best I can tell)
The bundled llama.cpp in this backend build only partially recognizes the
gemma4 GGUF architecture — it maps 720 of the declared 2131 tensors before
giving up, rather than either fully supporting the architecture or cleanly
rejecting it up front. This tracks with upstream ggml-org/llama.cpp gemma4
support still being in active flux — several open issues/PRs there reference
gemma4-family loading/eval problems, e.g.:
- https://github.com/ggml-org/llama.cpp/issues/25668 ("Eval bug: Can't run
gemma4-style model") - https://github.com/ggml-org/llama.cpp/pull/25439 ("model: Gemma4 12B model
type detection", still open) - https://github.com/ggml-org/llama.cpp/pull/25549 ("spec: support DSpark for
Gemma4", still open)
None of those match this exact "wrong number of tensors; expected 2131, got
720" symptom, so this looks like a distinct manifestation of the same
underlying "gemma4 architecture support isn't finished in llama.cpp yet" gap,
surfaced through LocalAI's bundled backend.
Ask
- If this is purely an upstream llama.cpp support gap, it'd help to have that
stated/tracked here (e.g. a pointer to which llama.cpp PR/milestone brings
fullgemma4tensor support) so downstream LocalAI users know when to
retry. - Separately: LocalAI/the backend should not let this crash the whole backend
process — a clean "unsupported architecture" load-time error (instead of
exitCode -1process death class of failure) would make this kind of gap
much easier to diagnose from the API side alone, without digging through
container logs.
Happy to provide the full GGUF metadata dump or test further builds if useful.
- 主要言語
- Go
- スター
- 49.2k
- フォーク
- 4.5k
- 平均マージ
- 1日 7時間
- マージ済み PR(30日)
- 340
環境構築
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
mudler/LocalAI のほかの issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
mudler/LocalAI#11995 · コメント 1 件 ·
メンテナーはふだん 3 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
mudler/LocalAI#11991 · コメント 1 件 ·
メンテナーはふだん 3 日以内に返信
-
fish-speech: make compile:true usable on Blackwell sm_121 by honouring the CUDA toolkit's ptxasオープンenhancement
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
mudler/LocalAI#11348 · コメント 1 件 ·
メンテナーはふだん 3 日以内に返信
-
feat: add automatic MCP transport selection for 2024-11-05 / 2025-03-26 / 2025-06-18 vs 2025-11-25オープンenhancement
難易度 3/5 1〜2日 初心者へのやさしさ 65/100
mudler/LocalAI#12262 · コメント 2 件 ·
メンテナーはふだん 3 日以内に返信
-
難易度 5/5 1週間以上 初心者へのやさしさ 25/100
メンテナーはふだん 3 日以内に返信
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
rossoctl/context-guru#346 ·
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 90/100
prime-radiant-inc/evener#2883 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
gravitational/teleport#69805 ·
メンテナーはふだん 11 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
メンテナーはふだん 1 日以内に返信
-
Under Poisson sampling, the `PLDAccountant` composes the inner event both before and after samplingオープン
難易度 2/5 半日 初心者へのやさしさ 78/100
google/differential-privacy#496 ·