Apple Silicon Metal + BLAS segfault for I2_S when ubatch >= 32 routes generic MUL_MAT

Đang mở Phù hợp với người mới
#512 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
2/5
Thời gian dự kiến
1-3 giờ
Mức phù hợp với người mới
72/100
Loại issue
Lỗi
Độ rõ ràng
Đặc tả rõ ràng
Mức độ hoạt động
Ít trao đổi
Công nghệ
cpp
Lĩnh vực
backend

Hướng nghiên cứu

Bắt đầu trong ggml-blas.cpp, tại phần kiểm tra hỗ trợ BLAS chung cho MUL_MAT, sau đó tái hiện vấn đề với Metal và Apple BLAS bằng cách sử dụng các giá trị i2_s và ubatch là 31 và 32 hoặc cao hơn. Xác nhận rằng guard I2_S giữ nguyên đường dẫn chuyên biệt và cấu hình với ubatch lớn hơn không còn gây ra segfault trong khi BLAS vẫn được bật.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Environment

  • Platform: Apple Silicon Mac
  • Host: Apple M4 Max
  • OS: macOS
  • Compiler: Homebrew clang 18.1.8
  • BitNet / submodule state: BitNet using vendored 3rdparty/llama.cpp at Eddie-Wang1120/llama.cpp commit 1f86f058de0c3f4098dedae2ae8653c335c868a1
  • Model: microsoft/BitNet-b1.58-2B-4T-gguf / ggml-model-i2_s.gguf
  • Build flags:
    • GGML_METAL=ON
    • GGML_ACCELERATE=ON
    • GGML_BLAS=ON
    • GGML_BLAS_VENDOR=Apple
    • BITNET_ARM_TL1=OFF

Problem

On Apple Silicon with Metal enabled, i2_s inference can segfault when BLAS is enabled and the physical micro-batch crosses the BLAS routing threshold.

The crash is tied to physical ubatch, not logical batch:

  • -b 2048 -ub 31 -> stable
  • -b 32 -ub 31 -> stable
  • -b 2048 -ub 32 -> segfault
  • -b 2048 -ub 512 -> segfault

This means the failure starts exactly when the BLAS backend begins claiming the generic MUL_MAT path for larger batches.

Control Experiment

The same Metal runtime is stable when BLAS is disabled:

  • BLAS ON + -b 2048 -ub 512 -> segfault
  • BLAS OFF + -b 2048 -ub 512 -> stable

This strongly suggests the crash is in the BLAS-side handling of GGML_TYPE_I2_S, not in Metal itself and not in the outer chat request schema.

Root Cause

ggml-blas.cpp allows the generic BLAS MUL_MAT path to accept quantized source tensors when ggml_get_type_traits(src0->type)->to_float != NULL.

For GGML_TYPE_I2_S, that is not safe:

  • I2_S stores an external scale outside the per-row payload
  • the generic BLAS dequantize-to-float path assumes self-contained per-row data
  • once ubatch >= 32, BLAS starts claiming MUL_MAT
  • that eventually crashes in the i2_s dequant / BLAS matmul path

In crash reports, the top frames consistently land in:

  • dequantize_row_i2_s
  • ggml_backend_blas_mul_mat

Proposed Fix

Reject GGML_TYPE_I2_S in the generic BLAS MUL_MAT support check so that I2_S continues using its specialized non-BLAS path:

return src0->type != GGML_TYPE_I2_S &&
       ggml_is_contiguous(src0) &&
       ggml_is_contiguous(src1) &&
       src1->type == GGML_TYPE_F32 &&
       (ne0 >= min_batch && ne1 >= min_batch && ne10 >= min_batch) &&
       (src0->type == GGML_TYPE_F32 || ggml_get_type_traits(src0->type)->to_float != NULL);

Result After Patch

After applying the BLAS guard above:

  • BLAS ON + Metal + -b 2048 -ub 512 is stable
  • managed broker end-to-end requests no longer segfault under the same settings

This does not solve all i2_s quality issues, but it does remove the native crash path.

Related Issues

  • #468
  • #470
  • #195
  • #411
Ngôn ngữ chính
C++
Star
40.3k
Fork
3.7k
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của microsoft/BitNet

Tất cả issue của microsoft/BitNet

Issue tương tự

Thêm issue về C++

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.