Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Supporting asynchronous HBM clock crossing and AXI width conversion

Aperta
#231 1 commento 0 reazioni 1 assegnatario Vedi su GitHub

@hpc-aulmamei ci sta già lavorando.

Dal 1/10/2026.

  • #232 di @ZhangMZh — aperta

Valutazione

Difficoltà
5/5
Tempo stimato
Più di una settimana
Idoneità per principianti
35/100
Tipo di issue
Funzionalità
Chiarezza
Abbastanza chiara
Stato di attività
Attiva

Direzione di ricerca

Start with examples/05_perf and inspect the draft implementation in commit 23d39e3f, focusing on the linker and generated vbin rather than the unchanged static shell. Reproduce the baseline and 512-bit, 200 MHz measurements, then validate the 1–64-channel scaling results. Done means the asynchronous clock crossing and AXI width conversion work upstream without requiring a shell rebuild.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Problem

The current HBM connection can limit bandwidth when the user kernel runs at a lower frequency than the static HBM AXI interface. Using 05_perf as an example, a 256-bit kernel at 200 MHz provides only 6.4 GB/s per channel, while the static shell at 360 MHz supports 11.52 GB/s per channel.

V80 provides approximately 819 GB/s of theoretical HBM bandwidth at 400 MHz. With the static shell’s 64 HBM AXI interfaces configured at 360 MHz, the aggregate interface ceiling is 737.28 GB/s. However, 05_perf achieves only the following results:

VRT Version: 1.0.0
Launching 64 perf kernels
Per-kernel buffer size: 512 MiB
Aggregate buffer footprint: 32.00 GiB
[2026-09-30 13:40:57.984] [INFO ] vrt::impl::Device::Device(const std::string&, const std::string&, bool, vrt::ProgramType): Programmed user clock to 193382352 Hz (target 193386192 Hz)
Write phase: launching 64 kernel(s)
Write phase time: 93 ms (368.30 GB/s aggregate)
Read phase: launching 64 kernel(s)
Read phase time: 87 ms (392.81 GB/s aggregate)
Combined read+write throughput: 380.16 GB/s aggregate
Test passed

Solution

The U280 User Guide recommends matching kernel bandwidth to the HBM interface through an appropriate combination of data width and clock frequency. For its 256-bit HBM interfaces at 450 MHz, the guide lists two kernel configurations:

  • 256 bits at 450 MHz
  • 512 bits at 225 MHz

The same principle can be applied to the V80 shell. Our draft implementation is available in commit 23d39e3f. We updated the linker to preserve the kernel’s AXI data width and generate a SmartConnect that performs asynchronous CDC and 512/256-bit width conversion. The static shell remains unchanged. No shell rebuild is required; only the application vbin needs to be regenerated. The kernel now uses a 512-bit interface targeting 200 MHz:

VRT Version: 1.0.0
Launching 64 perf kernels (perf_0 to perf_63)
Per-kernel buffer size: 512 MiB
Aggregate buffer footprint: 32.00 GiB
[2026-09-30 14:24:03.422] [INFO ] vrt::impl::Device::Device(const std::string&, const std::string&, bool, vrt::ProgramType): Programmed user clock to 197777777 Hz (target 197784810 Hz)
Write phase: launching 64 kernel(s)
Write phase time: 52 ms (654.08 GB/s aggregate)
Read phase: launching 64 kernel(s)
Read phase time: 48 ms (706.13 GB/s aggregate)
Combined read+write throughput: 681.40 GiB/s aggregate
Test passed

Scaling observations

We tested the modified 05_perf on V80 with 1–64 HBM channels, doubling the channel count between runs. With only a few channels, throughput closely matches the expected interface bandwidth. Aggregate throughput continues to increase with channel count, but efficiency decreases, particularly for writes. At 64 channels, read throughput reaches 95.8% of the interface ceiling, while write throughput reaches 88.7%.

Image

We have not yet investigated the NoC design in depth, so the cause of this scaling loss remains unclear. Could it involve bandwidth constraints or contention within the NoC/HBM subsystem, memory-controller behavior, or other effects?

The solution is an RM-side workaround that leaves the static shell unchanged. If this feature would be useful upstream, we can open a pull request and help adapt the implementation for integration.

Lingua principale
Tcl
Stelle
48
Fork
23
Merge medio
1g 22h
PR unite (30g)
6

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di Xilinx/SLASH

Tutte le issue di Xilinx/SLASH

Issue simili

Altre issue su Build System

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.