Model load holds two full copies of the weights on Metal (2.93 GB peak for a 1.42 GB GGUF)
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 48/100
- Tipo de issue
- Error
- Claridad
- Bastante claro
- Estado de actividad
- Tranquilo
- Stack tecnológico
- cpp, macos
- Área
- backend, performance
Línea de trabajo
Comienza siguiendo ModelLoader::load() y ModelLoader::realize_weights(), especialmente gguf_init_from_file(), el bucle de carga al dispositivo y la ruta del búfer de CPU. Compara el enfoque de carga y mmap en src/llama-model-loader.cpp de llama.cpp y, después, mide la memoria máxima y la memoria posterior a la carga en Metal. Se considera terminado cuando la carga al dispositivo ya no conserva una copia de staging anónima completa en el host, mientras las cargas siguen siendo correctas.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Summary
On Metal, loading a model transiently holds two full copies of the weights, making peak resident memory approximately twice the GGUF size. The generic non-CPU loading path appears to perform the same staging copy for CUDA and Vulkan, although I have only measured Metal.
Measured on an M2 with ctc-1.1b-q8_0.gguf:
| GGUF on disk | 1.42 GB |
| peak resident during load | 2.93 GB |
| resident after load | 1.51 GB |
The steady state is fine — it's the transient that doubles.
Where it comes from
ModelLoader::load() opens the file with no_alloc=false, so ggml allocates and reads every tensor into ordinary host memory:
struct gguf_init_params p{ /*no_alloc*/false, /*ctx*/&ctx_ };
gguf_ = gguf_init_from_file(path.c_str(), p);
ModelLoader::realize_weights() then, on the device path, allocates a second model-sized backend buffer and uploads each tensor from that host staging copy, releasing it only after every upload completes:
weights_buf_ = ggml_backend_alloc_ctx_tensors(device_ctx_, backend);
for (auto& pr : ups)
ggml_backend_tensor_set(pr.first, pr.second, 0, ggml_nbytes(pr.first));
...
ggml_free(ctx_);
Both copies are live for the duration of the upload loop, which accounts for the measured 1.42 GB → 2.93 GB → 1.51 GB progression: two weight copies plus metadata and alignment overhead.
The CPU path in the same function already avoids this, borrowing the loaded memory directly:
weights_buf_ = ggml_backend_cpu_buffer_from_ptr(base, size);
Why it matters
For an offline captioning app shipped to end users, the launch spike is what decides whether the app is usable on a small machine. On an 8 GB Mac, a 2.93 GB transient against a ~3-4 GB OS baseline creates substantial additional memory pressure and may cause compression or swapping depending on what else is running and on the model's compute buffers; 1.4 GB would leave considerably more headroom.
There's a second, smaller benefit: weights loaded this way are anonymous memory, so under pressure the OS can only compress or swap them. Weights backed by a file mapping can simply be dropped and re-read.
Possible direction
A broadly applicable first improvement may be to parse with no_alloc=true, mmap the GGUF tensor data, and upload tensors individually from that mapping. That would retain only the destination backend allocation plus file-backed source pages, avoiding the complete anonymous host staging allocation — a benefit on CUDA and Vulkan too, which still need backend-owned storage and a transfer.
On Apple Silicon specifically, ggml_backend_metal_buffer_from_ptr or the equivalent current Metal buffer API may permit the mapping itself to back the Metal weight buffers, potentially eliminating the copy entirely. llama.cpp takes a comparable approach — it parses with no_alloc=true and manages model storage separately, with mmap as a supported loading mode (use_mmap in src/llama-model-loader.cpp).
I haven't attempted a patch — I don't know whether the loader's tensor bookkeeping makes either step straightforward here, and you'd know immediately whether it's worth pursuing. Happy to test any change on the setup below.
Environment
- macOS 26.3.1, Apple M2, 24 GB
- parakeet.cpp v0.5.0 release build, Metal
ctc-1.1b-q8_0.gguffrommudler/parakeet-cpp-gguf
- Lenguaje dominante
- C++
- Estrellas
- 786
- Forks
- 93
- Merge medio
- 9 d 19 h
- PR fusionados (30 d)
- 4
Preparar el entorno
Aún no hemos revisado los archivos de configuración de este proyecto. Empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de mudler/parakeet.cpp
-
Dificultad 4/5 3-5 días Aptitud para principiantes 55/100
mudler/parakeet.cpp#68 ·
-
Dificultad 3/5 1-2 días Aptitud para principiantes 58/100
mudler/parakeet.cpp#62 ·
-
Dificultad 4/5 3-5 días Aptitud para principiantes 58/100
mudler/parakeet.cpp#61 ·
-
Real streaming from a micAbierto
Dificultad 3/5 1-2 días Aptitud para principiantes 48/100
mudler/parakeet.cpp#60 ·
-
Dificultad 4/5 3-5 días Aptitud para principiantes 45/100
mudler/parakeet.cpp#55 · 1 comentario ·
Todos los issues de mudler/parakeet.cpp
Issues similares
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 92/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
cp-algorithms/cp-algorithms#1715 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Icinga/icinga2#11058 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
status:needs-triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
PX4/PX4-Autopilot#28924 ·
Los mantenedores suelen responder en 1 día
-
component: split-view platform: windows
Dificultad 2/5 1-3 horas Aptitud para principiantes 74/100
zen-browser/desktop#15616 · 1 reacción ·
Los mantenedores suelen responder en 1 día