Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Model load holds two full copies of the weights on Metal (2.93 GB peak for a 1.42 GB GGUF)

Đang mở
#59 3 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức phù hợp với người mới
48/100
Loại issue
Lỗi
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Ít trao đổi
Công nghệ
cpp, macos
Lĩnh vực
backend, performance

Hướng nghiên cứu

Bắt đầu bằng cách theo dõi ModelLoader::load() và ModelLoader::realize_weights(), đặc biệt là gguf_init_from_file(), vòng lặp tải lên thiết bị và đường dẫn bộ đệm CPU. So sánh cách tiếp cận tải và mmap trong src/llama-model-loader.cpp của llama.cpp, sau đó đo bộ nhớ đỉnh và bộ nhớ sau khi tải trên Metal. Hoàn tất có nghĩa là việc tải lên thiết bị không còn giữ lại một bản sao staging ẩn danh đầy đủ trên host, trong khi các lần tải lên vẫn chính xác.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Summary

On Metal, loading a model transiently holds two full copies of the weights, making peak resident memory approximately twice the GGUF size. The generic non-CPU loading path appears to perform the same staging copy for CUDA and Vulkan, although I have only measured Metal.

Measured on an M2 with ctc-1.1b-q8_0.gguf:

GGUF on disk 1.42 GB
peak resident during load 2.93 GB
resident after load 1.51 GB

The steady state is fine — it's the transient that doubles.

Where it comes from

ModelLoader::load() opens the file with no_alloc=false, so ggml allocates and reads every tensor into ordinary host memory:

struct gguf_init_params p{ /*no_alloc*/false, /*ctx*/&ctx_ };
gguf_ = gguf_init_from_file(path.c_str(), p);

ModelLoader::realize_weights() then, on the device path, allocates a second model-sized backend buffer and uploads each tensor from that host staging copy, releasing it only after every upload completes:

weights_buf_ = ggml_backend_alloc_ctx_tensors(device_ctx_, backend);
for (auto& pr : ups)
    ggml_backend_tensor_set(pr.first, pr.second, 0, ggml_nbytes(pr.first));
...
ggml_free(ctx_);

Both copies are live for the duration of the upload loop, which accounts for the measured 1.42 GB → 2.93 GB → 1.51 GB progression: two weight copies plus metadata and alignment overhead.

The CPU path in the same function already avoids this, borrowing the loaded memory directly:

weights_buf_ = ggml_backend_cpu_buffer_from_ptr(base, size);

Why it matters

For an offline captioning app shipped to end users, the launch spike is what decides whether the app is usable on a small machine. On an 8 GB Mac, a 2.93 GB transient against a ~3-4 GB OS baseline creates substantial additional memory pressure and may cause compression or swapping depending on what else is running and on the model's compute buffers; 1.4 GB would leave considerably more headroom.

There's a second, smaller benefit: weights loaded this way are anonymous memory, so under pressure the OS can only compress or swap them. Weights backed by a file mapping can simply be dropped and re-read.

Possible direction

A broadly applicable first improvement may be to parse with no_alloc=true, mmap the GGUF tensor data, and upload tensors individually from that mapping. That would retain only the destination backend allocation plus file-backed source pages, avoiding the complete anonymous host staging allocation — a benefit on CUDA and Vulkan too, which still need backend-owned storage and a transfer.

On Apple Silicon specifically, ggml_backend_metal_buffer_from_ptr or the equivalent current Metal buffer API may permit the mapping itself to back the Metal weight buffers, potentially eliminating the copy entirely. llama.cpp takes a comparable approach — it parses with no_alloc=true and manages model storage separately, with mmap as a supported loading mode (use_mmap in src/llama-model-loader.cpp).

I haven't attempted a patch — I don't know whether the loader's tensor bookkeeping makes either step straightforward here, and you'd know immediately whether it's worth pursuing. Happy to test any change on the setup below.

Environment

  • macOS 26.3.1, Apple M2, 24 GB
  • parakeet.cpp v0.5.0 release build, Metal
  • ctc-1.1b-q8_0.gguf from mudler/parakeet-cpp-gguf
Ngôn ngữ chính
C++
Star
786
Fork
93
Merge trung bình
9 ngày 19 giờ
Pull request đã merge (30 ngày)
4

Chuẩn bị môi trường

Chúng tôi chưa kiểm tra các tệp thiết lập môi trường của dự án này. Hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của mudler/parakeet.cpp

Tất cả issue của mudler/parakeet.cpp

Issue tương tự

Thêm issue về C++

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.