cp.async ergonomics: missing address-space-aware intrinsic forces inline PTX
まだ誰も着手していません。
評価
調査の方向性
まず、Rust-CUDA の intrinsic または helper の定義と、cp.async に対する既存の PTX 処理を探します。この issue では具体的なファイルやテストは指定されていません。完了条件は、アドレス空間を考慮した unsafe な cp.async.cg.shared.global API が、ユーザーにアドレス変換や繰り返しの inline PTX の記述を要求せずに、説明された 16 バイトのパスをサポートすることです。
索引モデルが issue の本文から書いたものです。
説明
Summary
Using cp.async from Rust-CUDA currently requires handwritten inline PTX plus manual shared-address conversion in user kernels.
In practice, the missing piece is an address-space-aware API/intrinsic for cp.async.cg.shared.global.
Problem
For an async copy path like:
- global source pointer
- shared-memory destination
- 16-byte copy (
4 x f32)
users currently end up writing something like:
unsafe fn shared_addr(ptr: *mut f32) -> u32 {
let mut addr: u64;
asm!(
"cvta.to.shared.u64 {dst}, {src};",
dst = out(reg64) addr,
src = in(reg64) ptr,
);
addr as u32
}
unsafe fn cp_async4(dst_shared: u32, src_global: *const f32) {
asm!(
"cp.async.cg.shared.global [{dst}], [{src}], 16;",
dst = in(reg32) dst_shared,
src = in(reg64) src_global,
);
}
and then precompute shared base addresses in the kernel and pass integer offsets into the helper.
If we instead pass generic/raw pointers through a helper, PTX tends to contain extra address-space conversion glue around the cp.async sites.
Why this matters
This is exactly the kind of operation where users want:
- a small, explicit intrinsic
- correct shared/global address-space handling
- no repeated manual inline PTX in every project
Right now, getting good code requires low-level PTX knowledge and manual control over shared address conversion.
Requested improvement
A Rust-CUDA intrinsic/helper for cp.async.cg.shared.global (or equivalent family) that:
- represents the destination as shared-memory address space explicitly
- avoids forcing users to manually convert
*mut Tinto au32shared address - maps cleanly to the PTX async-copy instructions used in modern CUDA kernels
Even a low-level unsafe API would be useful if it preserves the right address-space semantics and avoids generic-pointer friction.
- 主要言語
- Rust
- スター
- 5.4k
- フォーク
- 249
- 平均マージ
- 5日 20時間
- マージ済み PR(30日)
- 2
環境構築
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
Rust-GPU/rust-cuda のほかの issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
-
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
-
Default NvvmArch::Compute75 silently produces InvalidPtx on pre-Turing GPUs (Pascal/Maxwell/Volta)オープン
難易度 4/5 3〜5日 初心者へのやさしさ 48/100
-
難易度 5/5 1週間以上 初心者へのやさしさ 30/100
-
難易度 4/5 3〜5日 初心者へのやさしさ 55/100
Rust-GPU/rust-cuda の issue をすべて見る
似ている issue
-
backend::vllm diffusion multimodal
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
lambdaclass/ethrex#7329 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
shadowsocks/shadowsocks-rust#2186 · コメント 1 件 ·
-
C-bug S-awaiting-triage
難易度 1/5 1時間未満 初心者へのやさしさ 92/100
juspay/hyperswitch#14479 ·
メンテナーはふだん 1 日以内に返信