Supporting asynchronous HBM clock crossing and AXI width conversion
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 35/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Ambito
- build-system, performance
Direzione di ricerca
Start with examples/05_perf and inspect the draft implementation in commit 23d39e3f, focusing on the linker and generated vbin rather than the unchanged static shell. Reproduce the baseline and 512-bit, 200 MHz measurements, then validate the 1–64-channel scaling results. Done means the asynchronous clock crossing and AXI width conversion work upstream without requiring a shell rebuild.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Problem
The current HBM connection can limit bandwidth when the user kernel runs at a lower frequency than the static HBM AXI interface. Using 05_perf as an example, a 256-bit kernel at 200 MHz provides only 6.4 GB/s per channel, while the static shell at 360 MHz supports 11.52 GB/s per channel.
V80 provides approximately 819 GB/s of theoretical HBM bandwidth at 400 MHz. With the static shell’s 64 HBM AXI interfaces configured at 360 MHz, the aggregate interface ceiling is 737.28 GB/s. However, 05_perf achieves only the following results:
VRT Version: 1.0.0
Launching 64 perf kernels
Per-kernel buffer size: 512 MiB
Aggregate buffer footprint: 32.00 GiB
[2026-09-30 13:40:57.984] [INFO ] vrt::impl::Device::Device(const std::string&, const std::string&, bool, vrt::ProgramType): Programmed user clock to 193382352 Hz (target 193386192 Hz)
Write phase: launching 64 kernel(s)
Write phase time: 93 ms (368.30 GB/s aggregate)
Read phase: launching 64 kernel(s)
Read phase time: 87 ms (392.81 GB/s aggregate)
Combined read+write throughput: 380.16 GB/s aggregate
Test passed
Solution
The U280 User Guide recommends matching kernel bandwidth to the HBM interface through an appropriate combination of data width and clock frequency. For its 256-bit HBM interfaces at 450 MHz, the guide lists two kernel configurations:
- 256 bits at 450 MHz
- 512 bits at 225 MHz
The same principle can be applied to the V80 shell. Our draft implementation is available in commit 23d39e3f. We updated the linker to preserve the kernel’s AXI data width and generate a SmartConnect that performs asynchronous CDC and 512/256-bit width conversion. The static shell remains unchanged. No shell rebuild is required; only the application vbin needs to be regenerated. The kernel now uses a 512-bit interface targeting 200 MHz:
VRT Version: 1.0.0
Launching 64 perf kernels (perf_0 to perf_63)
Per-kernel buffer size: 512 MiB
Aggregate buffer footprint: 32.00 GiB
[2026-09-30 14:24:03.422] [INFO ] vrt::impl::Device::Device(const std::string&, const std::string&, bool, vrt::ProgramType): Programmed user clock to 197777777 Hz (target 197784810 Hz)
Write phase: launching 64 kernel(s)
Write phase time: 52 ms (654.08 GB/s aggregate)
Read phase: launching 64 kernel(s)
Read phase time: 48 ms (706.13 GB/s aggregate)
Combined read+write throughput: 681.40 GiB/s aggregate
Test passed
Scaling observations
We tested the modified 05_perf on V80 with 1–64 HBM channels, doubling the channel count between runs. With only a few channels, throughput closely matches the expected interface bandwidth. Aggregate throughput continues to increase with channel count, but efficiency decreases, particularly for writes. At 64 channels, read throughput reaches 95.8% of the interface ceiling, while write throughput reaches 88.7%.
We have not yet investigated the NoC design in depth, so the cause of this scaling loss remains unclear. Could it involve bandwidth constraints or contention within the NoC/HBM subsystem, memory-controller behavior, or other effects?
The solution is an RM-side workaround that leaves the static shell unchanged. If this feature would be useful upstream, we can open a pull request and help adapt the implementation for integration.
- Lingua principale
- Tcl
- Stelle
- 48
- Fork
- 23
- Merge medio
- 1g 22h
- PR unite (30g)
- 6
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Nessun modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di Xilinx/SLASH
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 35/100
-
device bring up issue (v80-smi write-static-shell)Forse già presa @quetric l’ha presa 9 giorni fa. Aperta
Difficoltà 4/5 3-5 giorni Idoneità per principianti 35/100
-
V80 SLASH shell silently drops host `s_axi_control` writes to offsets `0x40–0xBF`Forse già presa @quetric l’ha presa 13 giorni fa. Aperta
Difficoltà 4/5 3-5 giorni Idoneità per principianti 45/100
-
The link is up and RX sees the packet, but the received packet is not counted as a good packet.Aperta
Difficoltà 4/5 3-5 giorni Idoneità per principianti 38/100
Tutte le issue di Xilinx/SLASH
Issue simili
-
Remove unused ts-node dependencyForse già presa @marcelofukumoto l’ha presa oggi. Apertaarea/dashboard kind/tech-debt QA/None
Difficoltà 2/5 1-3 ore Idoneità per principianti 83/100
I maintainer di solito rispondono entro 3 giorni
-
electron tech debt
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 66/100
Algorithmiq/monoprop#390 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
callstack/agent-device#3296 ·
I maintainer di solito rispondono entro 1 giorno
-
good first issue help wanted
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100