Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

cp.async ergonomics: missing address-space-aware intrinsic forces inline PTX

Đang mở
#378 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
38/100
Loại issue
Tính năng
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Ít trao đổi
Công nghệ
rust
Lĩnh vực
tooling

Hướng nghiên cứu

Bắt đầu bằng cách tìm các định nghĩa intrinsic hoặc helper của Rust-CUDA và phần xử lý PTX hiện có cho cp.async; issue không nêu các tệp hoặc test cụ thể. Công việc được xem là hoàn tất khi một API unsafe cp.async.cg.shared.global có nhận biết không gian địa chỉ hỗ trợ đường đi 16 byte được mô tả mà không yêu cầu người dùng tự viết phần chuyển đổi địa chỉ hoặc inline PTX lặp lại.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Summary

Using cp.async from Rust-CUDA currently requires handwritten inline PTX plus manual shared-address conversion in user kernels.

In practice, the missing piece is an address-space-aware API/intrinsic for cp.async.cg.shared.global.

Problem

For an async copy path like:

  • global source pointer
  • shared-memory destination
  • 16-byte copy (4 x f32)

users currently end up writing something like:

unsafe fn shared_addr(ptr: *mut f32) -> u32 {
    let mut addr: u64;
    asm!(
        "cvta.to.shared.u64 {dst}, {src};",
        dst = out(reg64) addr,
        src = in(reg64) ptr,
    );
    addr as u32
}

unsafe fn cp_async4(dst_shared: u32, src_global: *const f32) {
    asm!(
        "cp.async.cg.shared.global [{dst}], [{src}], 16;",
        dst = in(reg32) dst_shared,
        src = in(reg64) src_global,
    );
}

and then precompute shared base addresses in the kernel and pass integer offsets into the helper.

If we instead pass generic/raw pointers through a helper, PTX tends to contain extra address-space conversion glue around the cp.async sites.

Why this matters

This is exactly the kind of operation where users want:

  • a small, explicit intrinsic
  • correct shared/global address-space handling
  • no repeated manual inline PTX in every project

Right now, getting good code requires low-level PTX knowledge and manual control over shared address conversion.

Requested improvement

A Rust-CUDA intrinsic/helper for cp.async.cg.shared.global (or equivalent family) that:

  • represents the destination as shared-memory address space explicitly
  • avoids forcing users to manually convert *mut T into a u32 shared address
  • maps cleanly to the PTX async-copy instructions used in modern CUDA kernels

Even a low-level unsafe API would be useful if it preserves the right address-space semantics and avoids generic-pointer friction.

Ngôn ngữ chính
Rust
Star
5.4k
Fork
251
Merge trung bình
4 ngày 15 giờ
Pull request đã merge (30 ngày)
4

Chuẩn bị môi trường

Mở trong Codespaces

Khởi chạy dev container của dự án ngay trên trình duyệt, bằng tài khoản GitHub của bạn.

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của Rust-GPU/rust-cuda

Tất cả issue của Rust-GPU/rust-cuda

Issue tương tự

Thêm issue về Rust

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.