Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

feat(BACKEND-ROCM-F16-WEIGHTS): retain F16 dense and embedding weights on ROCm

Đang mở
#3,092 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
28/100
Loại issue
Tính năng
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
cpp
Lĩnh vực
backend, performance

Hướng nghiên cứu

Bắt đầu bằng việc đọc .agents/specs/rocm-f16-weights.md một lần sau khi tệp này được commit, sau đó tham khảo các nguồn upstream đã được ghim để biết hợp đồng số học cho đầu vào hỗn hợp. Theo dõi đường dẫn tải và sinh mặc định cho các trọng số dense và embedding, đồng thời giữ các expert xếp chồng trên đường dẫn mở rộng hiện có của chúng. Được xem là hoàn tất khi có các kiểm thử red-first, bằng chứng về tính đúng đắn và bộ nhớ trên gfx1100, việc rà soát mutation và xác minh của operator.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Row: BACKEND-ROCM-F16-WEIGHTS

ROCm currently expands GGUF F16 weights because ordinary matrix multiplication and embedding kernels reject F16 storage. This prevents those weights from staying in their checkpoint format on gfx1100.

Implement F16 storage support through ordinary Matmul/MatmulBT and embedding operations, preserving the existing BF16/F32 activation and output contract. Enable loader admission for dense matrix weights and embedding tables only. Keep stacked expert weights on their existing expansion path; this issue does not enable a full F16 activation runtime.

The spec must establish mixed input arithmetic from the pinned upstream sources, cover both ID widths for embeddings, and prove the default public loading and generation path reaches the new providers. Require red-first tests, physical gfx1100 correctness and memory evidence, fresh mutation review, and operator verification.

This work is independent of PR #2782. That PR expands quantized GEMM kernels. F16 RMSNorm activation support remains a distinct concern tracked by #2542. No CI changes are included.

Spec: .agents/specs/rocm-f16-weights.md (to be committed before implementation).

Ngôn ngữ chính
C++
Star
440
Fork
55
Merge trung bình
1 ngày 4 giờ
Pull request đã merge (30 ngày)
346

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của mudler/vllm.cpp

Tất cả issue của mudler/vllm.cpp

Issue tương tự

Thêm issue về C++

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.