Model load holds two full copies of the weights on Metal (2.93 GB peak for a 1.42 GB GGUF)
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 48/100
- Issue 类型
- 缺陷
- 描述清晰度
- 基本清楚
- 活跃度
- 冷清
- 技术栈
- cpp, macos
- 领域
- backend, performance
调研方向
首先跟踪 ModelLoader::load() 和 ModelLoader::realize_weights(),尤其是 gguf_init_from_file()、设备上传循环以及 CPU 缓冲区路径。对比 llama.cpp 的 src/llama-model-loader.cpp 中的加载和 mmap 方式,然后测量 Metal 上的峰值内存和加载后的内存。完成的标准是,设备加载不再保留完整的匿名主机暂存副本,同时上传仍然正确。
由索引模型根据 Issue 内容生成。
描述
Summary
On Metal, loading a model transiently holds two full copies of the weights, making peak resident memory approximately twice the GGUF size. The generic non-CPU loading path appears to perform the same staging copy for CUDA and Vulkan, although I have only measured Metal.
Measured on an M2 with ctc-1.1b-q8_0.gguf:
| GGUF on disk | 1.42 GB |
| peak resident during load | 2.93 GB |
| resident after load | 1.51 GB |
The steady state is fine — it's the transient that doubles.
Where it comes from
ModelLoader::load() opens the file with no_alloc=false, so ggml allocates and reads every tensor into ordinary host memory:
struct gguf_init_params p{ /*no_alloc*/false, /*ctx*/&ctx_ };
gguf_ = gguf_init_from_file(path.c_str(), p);
ModelLoader::realize_weights() then, on the device path, allocates a second model-sized backend buffer and uploads each tensor from that host staging copy, releasing it only after every upload completes:
weights_buf_ = ggml_backend_alloc_ctx_tensors(device_ctx_, backend);
for (auto& pr : ups)
ggml_backend_tensor_set(pr.first, pr.second, 0, ggml_nbytes(pr.first));
...
ggml_free(ctx_);
Both copies are live for the duration of the upload loop, which accounts for the measured 1.42 GB → 2.93 GB → 1.51 GB progression: two weight copies plus metadata and alignment overhead.
The CPU path in the same function already avoids this, borrowing the loaded memory directly:
weights_buf_ = ggml_backend_cpu_buffer_from_ptr(base, size);
Why it matters
For an offline captioning app shipped to end users, the launch spike is what decides whether the app is usable on a small machine. On an 8 GB Mac, a 2.93 GB transient against a ~3-4 GB OS baseline creates substantial additional memory pressure and may cause compression or swapping depending on what else is running and on the model's compute buffers; 1.4 GB would leave considerably more headroom.
There's a second, smaller benefit: weights loaded this way are anonymous memory, so under pressure the OS can only compress or swap them. Weights backed by a file mapping can simply be dropped and re-read.
Possible direction
A broadly applicable first improvement may be to parse with no_alloc=true, mmap the GGUF tensor data, and upload tensors individually from that mapping. That would retain only the destination backend allocation plus file-backed source pages, avoiding the complete anonymous host staging allocation — a benefit on CUDA and Vulkan too, which still need backend-owned storage and a transfer.
On Apple Silicon specifically, ggml_backend_metal_buffer_from_ptr or the equivalent current Metal buffer API may permit the mapping itself to back the Metal weight buffers, potentially eliminating the copy entirely. llama.cpp takes a comparable approach — it parses with no_alloc=true and manages model storage separately, with mmap as a supported loading mode (use_mmap in src/llama-model-loader.cpp).
I haven't attempted a patch — I don't know whether the loader's tensor bookkeeping makes either step straightforward here, and you'd know immediately whether it's worth pursuing. Happy to test any change on the setup below.
Environment
- macOS 26.3.1, Apple M2, 24 GB
- parakeet.cpp v0.5.0 release build, Metal
ctc-1.1b-q8_0.gguffrommudler/parakeet-cpp-gguf
- 主要语言
- C++
- 星标
- 786
- 派生
- 93
- 平均合并
- 9 天 19 小时
- 30 天内合并 PR
- 4
环境准备
我们还没有检查这个项目的环境配置文件。先看它的 README,通用步骤见我们的新手贡献指南。
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
mudler/parakeet.cpp 的其他 Issue
-
难度 4/5 3-5 天 新手友好度 55/100
mudler/parakeet.cpp#68 ·
-
难度 3/5 1-2 天 新手友好度 58/100
mudler/parakeet.cpp#62 ·
-
难度 4/5 3-5 天 新手友好度 58/100
mudler/parakeet.cpp#61 ·
-
难度 3/5 1-2 天 新手友好度 48/100
mudler/parakeet.cpp#60 ·
-
难度 4/5 3-5 天 新手友好度 45/100
mudler/parakeet.cpp#55 · 1 条评论 ·
查看 mudler/parakeet.cpp 的全部 Issue
相似的 Issue
-
难度 1/5 1 小时以内 新手友好度 92/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 72/100
cp-algorithms/cp-algorithms#1715 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 78/100
Icinga/icinga2#11058 · 1 条评论 ·
维护者通常 1 天内回复
-
status:needs-triage
难度 2/5 1-3 小时 新手友好度 88/100
PX4/PX4-Autopilot#28924 ·
维护者通常 1 天内回复
-
component: split-view platform: windows
难度 2/5 1-3 小时 新手友好度 74/100
zen-browser/desktop#15616 · 1 个 reaction ·
维护者通常 1 天内回复