/api/ps and /api/tags report size and size_vram as a hardcoded 0
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
調査の方向性
issue で参照されている /api/ps と /api/tags のマッピングについて、core/http/endpoints/ollama/models.go から始めます。ロード済みの llama-cpp モデルがどのように表現されているか、また resident size が利用可能かを追跡し、その後、ロード済みモデルを使って両方のエンドポイントを検証します。完了条件は、レスポンスで利用できないサイズ値が信頼できるゼロとして提示されなくなり、バックエンドが報告できる場合には実際のバイト数が保持されることです。
索引モデルが issue の本文から書いたものです。
説明
LocalAI version:
v4.9.0 (f7ad3f70eb5d8a0ddf80e08557f0d7df28cf032e), image localai/localai:v4.9.0-gpu-nvidia-cuda-12
Environment, CPU architecture, OS, and Version:
x86_64, RTX 3060 12GB, Docker on WSL2Describe the bug
GET /api/ps reports "size": 0 and "size_vram": 0 for every loaded model. The
values are literals rather than unpopulated fields:
https://github.com/mudler/LocalAI/blob/v4.9.0/core/http/endpoints/ollama/models.go#L106-L110
GET /api/tags has the same literal Size: 0 at line 40, and ExpiresAt on the ps
entry is synthesised as time.Now().Add(24 * time.Hour).
This is worse than omitting the fields. /api/ps is an Ollama-compatibility surface,
and Ollama reports real byte counts there, so a client written against Ollama reads
0 as authoritative and concludes the models cost nothing. Anything scheduling GPU
work from that will over-commit the card. An absent field or a null is safe, because
a consumer can distinguish "unknown" from "nothing"; a 0 cannot be distinguished.
To Reproduce
- Load any model (
granite-4.1-8b, llama-cpp backend) so it is resident. curl localhost:8080/api/ps- Compare against
nvidia-smi.
Expected behavior
Either the real resident sizes, or the fields omitted/null when the backend cannot
report them. Not 0.
Logs
{"models":[{"name":"granite-4.1-8b:latest","size":0,"size_vram":0,
"details":{"format":"gguf","family":"llama-cpp","parameter_size":"8B",
"quantization_level":"Q4_K_M"}}]}
Card at the same moment: 10849 / 12288 MiB used, two models resident (second row
elided, also all-zero).
Additional context
I gave up on /api/ps for this engine entirely and treat the size as unknown unconditionally, because I can't tell its zeros from real ones.
Filed separately as a feature request: there is no endpoint that reports per-model memory at all.
I'm not in a position to take the PR, but I'm happy to test a fix against this setup.
Written by my beloved Claude Code
- 主要言語
- Go
- スター
- 49.2k
- フォーク
- 4.5k
- 平均マージ
- 1日 7時間
- マージ済み PR(30日)
- 357
環境構築
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
mudler/LocalAI のほかの issue
-
bug unconfirmed
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
mudler/LocalAI#12337 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
mudler/LocalAI#11995 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
mudler/LocalAI#11991 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
fish-speech: make compile:true usable on Blackwell sm_121 by honouring the CUDA toolkit's ptxasオープンenhancement
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
mudler/LocalAI#11348 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
bug unconfirmed
難易度 4/5 3〜5日 初心者へのやさしさ 70/100
mudler/LocalAI#12331 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
似ている issue
-
bug needs-triage
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
DataDog/dd-trace-go#5469 ·
メンテナーはふだん 1 日以内に返信
-
bug tests
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
l3montree-dev/devguard#3101 ·
メンテナーはふだん 1 日以内に返信
-
area:*of bug
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
oapi-codegen/oapi-codegen#2593 ·
メンテナーはふだん 1 日以内に返信
-
bug
難易度 1/5 1時間未満 初心者へのやさしさ 85/100
DaoCloud/DaoCloud-docs#7432 ·
メンテナーはふだん 1 日以内に返信