Supporting asynchronous HBM clock crossing and AXI width conversion
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 35/100
- Loại issue
- Tính năng
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Lĩnh vực
- build-system, performance
Hướng nghiên cứu
Start with examples/05_perf and inspect the draft implementation in commit 23d39e3f, focusing on the linker and generated vbin rather than the unchanged static shell. Reproduce the baseline and 512-bit, 200 MHz measurements, then validate the 1–64-channel scaling results. Done means the asynchronous clock crossing and AXI width conversion work upstream without requiring a shell rebuild.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Problem
The current HBM connection can limit bandwidth when the user kernel runs at a lower frequency than the static HBM AXI interface. Using 05_perf as an example, a 256-bit kernel at 200 MHz provides only 6.4 GB/s per channel, while the static shell at 360 MHz supports 11.52 GB/s per channel.
V80 provides approximately 819 GB/s of theoretical HBM bandwidth at 400 MHz. With the static shell’s 64 HBM AXI interfaces configured at 360 MHz, the aggregate interface ceiling is 737.28 GB/s. However, 05_perf achieves only the following results:
VRT Version: 1.0.0
Launching 64 perf kernels
Per-kernel buffer size: 512 MiB
Aggregate buffer footprint: 32.00 GiB
[2026-09-30 13:40:57.984] [INFO ] vrt::impl::Device::Device(const std::string&, const std::string&, bool, vrt::ProgramType): Programmed user clock to 193382352 Hz (target 193386192 Hz)
Write phase: launching 64 kernel(s)
Write phase time: 93 ms (368.30 GB/s aggregate)
Read phase: launching 64 kernel(s)
Read phase time: 87 ms (392.81 GB/s aggregate)
Combined read+write throughput: 380.16 GB/s aggregate
Test passed
Solution
The U280 User Guide recommends matching kernel bandwidth to the HBM interface through an appropriate combination of data width and clock frequency. For its 256-bit HBM interfaces at 450 MHz, the guide lists two kernel configurations:
- 256 bits at 450 MHz
- 512 bits at 225 MHz
The same principle can be applied to the V80 shell. Our draft implementation is available in commit 23d39e3f. We updated the linker to preserve the kernel’s AXI data width and generate a SmartConnect that performs asynchronous CDC and 512/256-bit width conversion. The static shell remains unchanged. No shell rebuild is required; only the application vbin needs to be regenerated. The kernel now uses a 512-bit interface targeting 200 MHz:
VRT Version: 1.0.0
Launching 64 perf kernels (perf_0 to perf_63)
Per-kernel buffer size: 512 MiB
Aggregate buffer footprint: 32.00 GiB
[2026-09-30 14:24:03.422] [INFO ] vrt::impl::Device::Device(const std::string&, const std::string&, bool, vrt::ProgramType): Programmed user clock to 197777777 Hz (target 197784810 Hz)
Write phase: launching 64 kernel(s)
Write phase time: 52 ms (654.08 GB/s aggregate)
Read phase: launching 64 kernel(s)
Read phase time: 48 ms (706.13 GB/s aggregate)
Combined read+write throughput: 681.40 GiB/s aggregate
Test passed
Scaling observations
We tested the modified 05_perf on V80 with 1–64 HBM channels, doubling the channel count between runs. With only a few channels, throughput closely matches the expected interface bandwidth. Aggregate throughput continues to increase with channel count, but efficiency decreases, particularly for writes. At 64 channels, read throughput reaches 95.8% of the interface ceiling, while write throughput reaches 88.7%.
We have not yet investigated the NoC design in depth, so the cause of this scaling loss remains unclear. Could it involve bandwidth constraints or contention within the NoC/HBM subsystem, memory-controller behavior, or other effects?
The solution is an RM-side workaround that leaves the static shell unchanged. If this feature would be useful upstream, we can open a pull request and help adapt the implementation for integration.
- Ngôn ngữ chính
- Tcl
- Star
- 48
- Fork
- 23
- Merge trung bình
- 1 ngày 22 giờ
- Pull request đã merge (30 ngày)
- 6
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Không có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của Xilinx/SLASH
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 35/100
-
device bring up issue (v80-smi write-static-shell)Có thể đã có người làm @quetric đã nhận 10 ngày trước. Đang mở
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 35/100
-
V80 SLASH shell silently drops host `s_axi_control` writes to offsets `0x40–0xBF`Có thể đã có người làm @quetric đã nhận 14 ngày trước. Đang mở
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 45/100
-
The link is up and RX sees the packet, but the received packet is not counted as a good packet.Đang mở
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 38/100
Issue tương tự
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 75/100
element-hq/lk-jwt-service#248 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
area/install-update comp/cli duplicate P2 python:uv sweeper:risk-compatibility type/bug
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 62/100
NousResearch/hermes-agent#135440 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
pytorch/tensordict#1927 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug ticket
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
cratestack/cratestack#1154 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
dependencies java
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 62/100
micrometer-metrics/tracing#1588 ·
Maintainer thường phản hồi trong vòng 1 ngày