Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

cp.async ergonomics: missing address-space-aware intrinsic forces inline PTX

オープン
#378 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
38/100
issue の種類
機能追加
明瞭さ
おおむね明確
活発さ
静か
技術スタック
rust
領域
tooling

調査の方向性

まず、Rust-CUDA の intrinsic または helper の定義と、cp.async に対する既存の PTX 処理を探します。この issue では具体的なファイルやテストは指定されていません。完了条件は、アドレス空間を考慮した unsafe な cp.async.cg.shared.global API が、ユーザーにアドレス変換や繰り返しの inline PTX の記述を要求せずに、説明された 16 バイトのパスをサポートすることです。

索引モデルが issue の本文から書いたものです。

説明

Summary

Using cp.async from Rust-CUDA currently requires handwritten inline PTX plus manual shared-address conversion in user kernels.

In practice, the missing piece is an address-space-aware API/intrinsic for cp.async.cg.shared.global.

Problem

For an async copy path like:

  • global source pointer
  • shared-memory destination
  • 16-byte copy (4 x f32)

users currently end up writing something like:

unsafe fn shared_addr(ptr: *mut f32) -> u32 {
    let mut addr: u64;
    asm!(
        "cvta.to.shared.u64 {dst}, {src};",
        dst = out(reg64) addr,
        src = in(reg64) ptr,
    );
    addr as u32
}

unsafe fn cp_async4(dst_shared: u32, src_global: *const f32) {
    asm!(
        "cp.async.cg.shared.global [{dst}], [{src}], 16;",
        dst = in(reg32) dst_shared,
        src = in(reg64) src_global,
    );
}

and then precompute shared base addresses in the kernel and pass integer offsets into the helper.

If we instead pass generic/raw pointers through a helper, PTX tends to contain extra address-space conversion glue around the cp.async sites.

Why this matters

This is exactly the kind of operation where users want:

  • a small, explicit intrinsic
  • correct shared/global address-space handling
  • no repeated manual inline PTX in every project

Right now, getting good code requires low-level PTX knowledge and manual control over shared address conversion.

Requested improvement

A Rust-CUDA intrinsic/helper for cp.async.cg.shared.global (or equivalent family) that:

  • represents the destination as shared-memory address space explicitly
  • avoids forcing users to manually convert *mut T into a u32 shared address
  • maps cleanly to the PTX async-copy instructions used in modern CUDA kernels

Even a low-level unsafe API would be useful if it preserves the right address-space semantics and avoids generic-pointer friction.

主要言語
Rust
スター
5.4k
フォーク
249
平均マージ
5日 20時間
マージ済み PR(30日)
2

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

Rust-GPU/rust-cuda のほかの issue

Rust-GPU/rust-cuda の issue をすべて見る

似ている issue

Rust の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。