[FEA]: Require ct.barrier for multi stage kernels
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 5/5
- Geschätzter Aufwand
- Über eine Woche
- Anfängerfreundlichkeit
- 35/100
Rechercherichtung
Beginne mit der Überprüfung der vorhandenen Einstiegspunkte ct.kernel, ct.load, ct.atomic_add und ct.store und ermittle anschließend, wie ein vorgeschlagenes ct.barrier Blöcke über mehrstufige Kernels hinweg koordinieren würde. Vergleiche die im Issue beschriebenen Ansätze mit einem Zähler im globalen Speicher und mit cooperative-groups. Die Aufgabe gilt als abgeschlossen, wenn eine dokumentierte Barrier-Funktion den beispielhaften Synchronisationsablauf unterstützt und eine Validierung ihrer Semantik vorhanden ist.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
Is this a new feature, an improvement, or a change to existing functionality?
New Feature
How would you describe the priority of this feature request?
High
Please provide a clear description of problem this feature solves
In CUDA programming, we use atomic methods or cooperative groups to synchronize execution across blocks.
cutile could provide a similar mechanism to help developers write complex multi-stage kernels in a simpler way.
Feature Description
Example:
import torch
import cuda.tile as ct
@ct.kernel
def device_norm(
x: ct.Array, y: ct.Array, workspace: ct.Array,
tile_size: ct.Constant, p: ct.Constant):
# create a barrier on global memory, except p blocks to reach it.
barrier = ct.barrier(p=p)
block_id = ct.bid(0)
tile = ct.load(x, index=(block_id, 0), shape=(1, tile_size))
mean = ct.sum(tile) / tile_size
ct.atomic_add(workspace, (0, ), mean)
# wait until p blocks to reach here
barrier.wait()
global_mean = ct.load(workspace, (0, ), (1, ))
global_mean = global_mean / p
tile = tile - global_mean
ct.store(y, (block_id, ), (tile_size, ))
Describe your ideal solution
Provide ct.barrier, or a similar feature, to make it easier for developers to write applications that require block-level synchronization.
There are multiple ways to implement ct.barrier:
- Allocate a region in global memory for synchronization, and let each block atomically increment a counter when it reaches the barrier.
- Use cooperative groups.
Describe any alternatives you have considered
No response
Additional context
No response
Contributing Guidelines
- I agree to follow cuTile Python's contributing guidelines
- I have searched the open feature requests and have found no duplicates for this feature request
- Vorherrschende Sprache
- Python
- Sterne
- 2.2k
- Forks
- 155
- PR-Merge-Kennzahlen
- Keine gemergten PRs in 30 T.
Entwicklungsumgebung
- Kein Dockerfile und keine Docker-Compose-Datei
- Hat eine Pull-Request-Vorlage
- Beitragsleitfaden lesen
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus NVIDIA/cutile-python
-
[BUG]: check_dtype_support rejects family-conditional (sm_XXXa) gpu_code targetsEvtl. vergeben @sylvesterkaczmarek hat das vor 31 Tagen übernommen. Offen
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 82/100
NVIDIA/cutile-python#105 · 2 Kommentare ·
-
nvidia-runners
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 25/100
NVIDIA/cutile-python#108 ·
-
[BUG]: FFT sample launches `Batch` blocks that each process the full batchEvtl. vergeben @AntonOresten hat das vor 138 Tagen übernommen. Offenbug status: needs-triage
Schwierigkeit 3/5 1-2 Tage Anfängerfreundlichkeit 68/100
NVIDIA/cutile-python#102 ·
-
Schwierigkeit 4/5 3-5 Tage Anfängerfreundlichkeit 68/100
NVIDIA/cutile-python#101 ·
-
bug
Schwierigkeit 4/5 3-5 Tage Anfängerfreundlichkeit 45/100
NVIDIA/cutile-python#97 · 1 Kommentar ·
Alle Issues in NVIDIA/cutile-python
Ähnliche Issues
-
namespace operations
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 72/100
EclipseFdn/open-vsx.org#14043 ·
Maintainer antworten meist innerhalb von 1 Tag
-
netbox status: needs triage type: bug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 76/100
netbox-community/netbox#23376 ·
Maintainer antworten meist innerhalb von 1 Tag
-
feedback simulation workshop
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 73/100
githubnext/gh-aw-workshop#4455 ·
Maintainer antworten meist innerhalb von 1 Tag
-
Triage 🩺
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 76/100
Maintainer antworten meist innerhalb von 1 Tag
-
[BUG] Container scenario crashes without expected_recovery_time, kube DNS example uses retry_waitOffenneeds-triage
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 77/100
krkn-chaos/krkn#1627 · 1 Kommentar ·
Maintainer antworten meist innerhalb von 1 Tag