Model load holds two full copies of the weights on Metal (2.93 GB peak for a 1.42 GB GGUF)
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 48/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Ít trao đổi
- Công nghệ
- cpp, macos
- Lĩnh vực
- backend, performance
Hướng nghiên cứu
Bắt đầu bằng cách theo dõi ModelLoader::load() và ModelLoader::realize_weights(), đặc biệt là gguf_init_from_file(), vòng lặp tải lên thiết bị và đường dẫn bộ đệm CPU. So sánh cách tiếp cận tải và mmap trong src/llama-model-loader.cpp của llama.cpp, sau đó đo bộ nhớ đỉnh và bộ nhớ sau khi tải trên Metal. Hoàn tất có nghĩa là việc tải lên thiết bị không còn giữ lại một bản sao staging ẩn danh đầy đủ trên host, trong khi các lần tải lên vẫn chính xác.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Summary
On Metal, loading a model transiently holds two full copies of the weights, making peak resident memory approximately twice the GGUF size. The generic non-CPU loading path appears to perform the same staging copy for CUDA and Vulkan, although I have only measured Metal.
Measured on an M2 with ctc-1.1b-q8_0.gguf:
| GGUF on disk | 1.42 GB |
| peak resident during load | 2.93 GB |
| resident after load | 1.51 GB |
The steady state is fine — it's the transient that doubles.
Where it comes from
ModelLoader::load() opens the file with no_alloc=false, so ggml allocates and reads every tensor into ordinary host memory:
struct gguf_init_params p{ /*no_alloc*/false, /*ctx*/&ctx_ };
gguf_ = gguf_init_from_file(path.c_str(), p);
ModelLoader::realize_weights() then, on the device path, allocates a second model-sized backend buffer and uploads each tensor from that host staging copy, releasing it only after every upload completes:
weights_buf_ = ggml_backend_alloc_ctx_tensors(device_ctx_, backend);
for (auto& pr : ups)
ggml_backend_tensor_set(pr.first, pr.second, 0, ggml_nbytes(pr.first));
...
ggml_free(ctx_);
Both copies are live for the duration of the upload loop, which accounts for the measured 1.42 GB → 2.93 GB → 1.51 GB progression: two weight copies plus metadata and alignment overhead.
The CPU path in the same function already avoids this, borrowing the loaded memory directly:
weights_buf_ = ggml_backend_cpu_buffer_from_ptr(base, size);
Why it matters
For an offline captioning app shipped to end users, the launch spike is what decides whether the app is usable on a small machine. On an 8 GB Mac, a 2.93 GB transient against a ~3-4 GB OS baseline creates substantial additional memory pressure and may cause compression or swapping depending on what else is running and on the model's compute buffers; 1.4 GB would leave considerably more headroom.
There's a second, smaller benefit: weights loaded this way are anonymous memory, so under pressure the OS can only compress or swap them. Weights backed by a file mapping can simply be dropped and re-read.
Possible direction
A broadly applicable first improvement may be to parse with no_alloc=true, mmap the GGUF tensor data, and upload tensors individually from that mapping. That would retain only the destination backend allocation plus file-backed source pages, avoiding the complete anonymous host staging allocation — a benefit on CUDA and Vulkan too, which still need backend-owned storage and a transfer.
On Apple Silicon specifically, ggml_backend_metal_buffer_from_ptr or the equivalent current Metal buffer API may permit the mapping itself to back the Metal weight buffers, potentially eliminating the copy entirely. llama.cpp takes a comparable approach — it parses with no_alloc=true and manages model storage separately, with mmap as a supported loading mode (use_mmap in src/llama-model-loader.cpp).
I haven't attempted a patch — I don't know whether the loader's tensor bookkeeping makes either step straightforward here, and you'd know immediately whether it's worth pursuing. Happy to test any change on the setup below.
Environment
- macOS 26.3.1, Apple M2, 24 GB
- parakeet.cpp v0.5.0 release build, Metal
ctc-1.1b-q8_0.gguffrommudler/parakeet-cpp-gguf
- Ngôn ngữ chính
- C++
- Star
- 786
- Fork
- 93
- Merge trung bình
- 9 ngày 19 giờ
- Pull request đã merge (30 ngày)
- 4
Chuẩn bị môi trường
Chúng tôi chưa kiểm tra các tệp thiết lập môi trường của dự án này. Hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của mudler/parakeet.cpp
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 55/100
mudler/parakeet.cpp#68 ·
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 58/100
mudler/parakeet.cpp#62 ·
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 58/100
mudler/parakeet.cpp#61 ·
-
Real streaming from a micĐang mở
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 48/100
mudler/parakeet.cpp#60 ·
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 45/100
mudler/parakeet.cpp#55 · 1 bình luận ·
Tất cả issue của mudler/parakeet.cpp
Issue tương tự
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
cp-algorithms/cp-algorithms#1715 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Icinga/icinga2#11058 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
status:needs-triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
PX4/PX4-Autopilot#28924 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
component: split-view platform: windows
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
zen-browser/desktop#15616 · 1 reaction ·
Maintainer thường phản hồi trong vòng 1 ngày