Supporting asynchronous HBM clock crossing and AXI width conversion
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 35/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Domain
- build-system, performance
Research direction
Start with examples/05_perf and inspect the draft implementation in commit 23d39e3f, focusing on the linker and generated vbin rather than the unchanged static shell. Reproduce the baseline and 512-bit, 200 MHz measurements, then validate the 1–64-channel scaling results. Done means the asynchronous clock crossing and AXI width conversion work upstream without requiring a shell rebuild.
Written by the indexing model from the issue text.
Description
Problem
The current HBM connection can limit bandwidth when the user kernel runs at a lower frequency than the static HBM AXI interface. Using 05_perf as an example, a 256-bit kernel at 200 MHz provides only 6.4 GB/s per channel, while the static shell at 360 MHz supports 11.52 GB/s per channel.
V80 provides approximately 819 GB/s of theoretical HBM bandwidth at 400 MHz. With the static shell’s 64 HBM AXI interfaces configured at 360 MHz, the aggregate interface ceiling is 737.28 GB/s. However, 05_perf achieves only the following results:
VRT Version: 1.0.0
Launching 64 perf kernels
Per-kernel buffer size: 512 MiB
Aggregate buffer footprint: 32.00 GiB
[2026-09-30 13:40:57.984] [INFO ] vrt::impl::Device::Device(const std::string&, const std::string&, bool, vrt::ProgramType): Programmed user clock to 193382352 Hz (target 193386192 Hz)
Write phase: launching 64 kernel(s)
Write phase time: 93 ms (368.30 GB/s aggregate)
Read phase: launching 64 kernel(s)
Read phase time: 87 ms (392.81 GB/s aggregate)
Combined read+write throughput: 380.16 GB/s aggregate
Test passed
Solution
The U280 User Guide recommends matching kernel bandwidth to the HBM interface through an appropriate combination of data width and clock frequency. For its 256-bit HBM interfaces at 450 MHz, the guide lists two kernel configurations:
- 256 bits at 450 MHz
- 512 bits at 225 MHz
The same principle can be applied to the V80 shell. Our draft implementation is available in commit 23d39e3f. We updated the linker to preserve the kernel’s AXI data width and generate a SmartConnect that performs asynchronous CDC and 512/256-bit width conversion. The static shell remains unchanged. No shell rebuild is required; only the application vbin needs to be regenerated. The kernel now uses a 512-bit interface targeting 200 MHz:
VRT Version: 1.0.0
Launching 64 perf kernels (perf_0 to perf_63)
Per-kernel buffer size: 512 MiB
Aggregate buffer footprint: 32.00 GiB
[2026-09-30 14:24:03.422] [INFO ] vrt::impl::Device::Device(const std::string&, const std::string&, bool, vrt::ProgramType): Programmed user clock to 197777777 Hz (target 197784810 Hz)
Write phase: launching 64 kernel(s)
Write phase time: 52 ms (654.08 GB/s aggregate)
Read phase: launching 64 kernel(s)
Read phase time: 48 ms (706.13 GB/s aggregate)
Combined read+write throughput: 681.40 GiB/s aggregate
Test passed
Scaling observations
We tested the modified 05_perf on V80 with 1–64 HBM channels, doubling the channel count between runs. With only a few channels, throughput closely matches the expected interface bandwidth. Aggregate throughput continues to increase with channel count, but efficiency decreases, particularly for writes. At 64 channels, read throughput reaches 95.8% of the interface ceiling, while write throughput reaches 88.7%.
We have not yet investigated the NoC design in depth, so the cause of this scaling loss remains unclear. Could it involve bandwidth constraints or contention within the NoC/HBM subsystem, memory-controller behavior, or other effects?
The solution is an RM-side workaround that leaves the static shell unchanged. If this feature would be useful upstream, we can open a pull request and help adapt the implementation for integration.
- Dominant language
- Tcl
- Stars
- 48
- Forks
- 23
- Avg merge
- 1d 22h
- Merged PRs (30d)
- 6
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Xilinx/SLASH
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
-
device bring up issue (v80-smi write-static-shell)Possibly taken @quetric claimed this 9 days ago. Open
Difficulty 4/5 3-5 days Newbie friendliness 35/100
-
V80 SLASH shell silently drops host `s_axi_control` writes to offsets `0x40–0xBF`Possibly taken @quetric claimed this 13 days ago. Open
Difficulty 4/5 3-5 days Newbie friendliness 45/100
-
The link is up and RX sees the packet, but the received packet is not counted as a good packet.Open
Difficulty 4/5 3-5 days Newbie friendliness 38/100
Similar issues
-
Remove unused ts-node dependencyPossibly taken @marcelofukumoto claimed this today. Openarea/dashboard kind/tech-debt QA/None
Difficulty 2/5 1-3 hours Newbie friendliness 83/100
Maintainers usually reply within 3 days
-
electron tech debt
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 66/100
Algorithmiq/monoprop#390 ·
Maintainers usually reply within 1 day
-
Area: Instruments Bug Difficulty: Low Priority: Low
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
DouglasNeuroInformatics/OpenDataCapture#1801 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
callstack/agent-device#3296 ·
Maintainers usually reply within 1 day