/api/ps and /api/tags report size and size_vram as a hardcoded 0
Maintainer thường phản hồi trong vòng 3 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 3/5
- Thời gian dự kiến
- 1-2 ngày
- Mức phù hợp với người mới
- 68/100
Hướng nghiên cứu
Bắt đầu trong core/http/endpoints/ollama/models.go, tại các ánh xạ /api/ps và /api/tags được issue tham chiếu. Theo dõi cách các mô hình llama-cpp đã được tải được biểu diễn và liệu kích thước resident có sẵn hay không; sau đó xác minh cả hai endpoint bằng một mô hình đã được tải. Hoàn thành có nghĩa là các phản hồi không còn trình bày các giá trị kích thước không khả dụng dưới dạng các số 0 có giá trị khẳng định, đồng thời vẫn giữ các số byte thực khi backend có thể báo cáo chúng.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
LocalAI version:
v4.9.0 (f7ad3f70eb5d8a0ddf80e08557f0d7df28cf032e), image localai/localai:v4.9.0-gpu-nvidia-cuda-12
Environment, CPU architecture, OS, and Version:
x86_64, RTX 3060 12GB, Docker on WSL2Describe the bug
GET /api/ps reports "size": 0 and "size_vram": 0 for every loaded model. The
values are literals rather than unpopulated fields:
https://github.com/mudler/LocalAI/blob/v4.9.0/core/http/endpoints/ollama/models.go#L106-L110
GET /api/tags has the same literal Size: 0 at line 40, and ExpiresAt on the ps
entry is synthesised as time.Now().Add(24 * time.Hour).
This is worse than omitting the fields. /api/ps is an Ollama-compatibility surface,
and Ollama reports real byte counts there, so a client written against Ollama reads
0 as authoritative and concludes the models cost nothing. Anything scheduling GPU
work from that will over-commit the card. An absent field or a null is safe, because
a consumer can distinguish "unknown" from "nothing"; a 0 cannot be distinguished.
To Reproduce
- Load any model (
granite-4.1-8b, llama-cpp backend) so it is resident. curl localhost:8080/api/ps- Compare against
nvidia-smi.
Expected behavior
Either the real resident sizes, or the fields omitted/null when the backend cannot
report them. Not 0.
Logs
{"models":[{"name":"granite-4.1-8b:latest","size":0,"size_vram":0,
"details":{"format":"gguf","family":"llama-cpp","parameter_size":"8B",
"quantization_level":"Q4_K_M"}}]}
Card at the same moment: 10849 / 12288 MiB used, two models resident (second row
elided, also all-zero).
Additional context
I gave up on /api/ps for this engine entirely and treat the size as unknown unconditionally, because I can't tell its zeros from real ones.
Filed separately as a feature request: there is no endpoint that reports per-model memory at all.
I'm not in a position to take the PR, but I'm happy to test a fix against this setup.
Written by my beloved Claude Code
- Ngôn ngữ chính
- Go
- Star
- 49.2k
- Fork
- 4.5k
- Merge trung bình
- 1 ngày 7 giờ
- Pull request đã merge (30 ngày)
- 340
Chuẩn bị môi trường
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của mudler/LocalAI
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
mudler/LocalAI#11995 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 3 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
mudler/LocalAI#11991 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 3 ngày
-
fish-speech: make compile:true usable on Blackwell sm_121 by honouring the CUDA toolkit's ptxasĐang mởenhancement
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
mudler/LocalAI#11348 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 3 ngày
-
feat: add automatic MCP transport selection for 2024-11-05 / 2025-03-26 / 2025-06-18 vs 2025-11-25Đang mởenhancement
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 65/100
mudler/LocalAI#12262 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 3 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 25/100
Maintainer thường phản hồi trong vòng 3 ngày
Tất cả issue của mudler/LocalAI
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
rossoctl/context-guru#346 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 90/100
prime-radiant-inc/evener#2883 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
gravitational/teleport#69805 ·
Maintainer thường phản hồi trong vòng 11 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Under Poisson sampling, the `PLDAccountant` composes the inner event both before and after samplingĐang mở
Độ khó 2/5 Nửa ngày Mức phù hợp với người mới 78/100
google/differential-privacy#496 ·