Supporting asynchronous HBM clock crossing and AXI width conversion
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 35/100
- Tipo de issue
- Nueva funcionalidad
- Claridad
- Bastante claro
- Estado de actividad
- Activo
- Área
- build-system, performance
Línea de trabajo
Start with examples/05_perf and inspect the draft implementation in commit 23d39e3f, focusing on the linker and generated vbin rather than the unchanged static shell. Reproduce the baseline and 512-bit, 200 MHz measurements, then validate the 1–64-channel scaling results. Done means the asynchronous clock crossing and AXI width conversion work upstream without requiring a shell rebuild.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Problem
The current HBM connection can limit bandwidth when the user kernel runs at a lower frequency than the static HBM AXI interface. Using 05_perf as an example, a 256-bit kernel at 200 MHz provides only 6.4 GB/s per channel, while the static shell at 360 MHz supports 11.52 GB/s per channel.
V80 provides approximately 819 GB/s of theoretical HBM bandwidth at 400 MHz. With the static shell’s 64 HBM AXI interfaces configured at 360 MHz, the aggregate interface ceiling is 737.28 GB/s. However, 05_perf achieves only the following results:
VRT Version: 1.0.0
Launching 64 perf kernels
Per-kernel buffer size: 512 MiB
Aggregate buffer footprint: 32.00 GiB
[2026-09-30 13:40:57.984] [INFO ] vrt::impl::Device::Device(const std::string&, const std::string&, bool, vrt::ProgramType): Programmed user clock to 193382352 Hz (target 193386192 Hz)
Write phase: launching 64 kernel(s)
Write phase time: 93 ms (368.30 GB/s aggregate)
Read phase: launching 64 kernel(s)
Read phase time: 87 ms (392.81 GB/s aggregate)
Combined read+write throughput: 380.16 GB/s aggregate
Test passed
Solution
The U280 User Guide recommends matching kernel bandwidth to the HBM interface through an appropriate combination of data width and clock frequency. For its 256-bit HBM interfaces at 450 MHz, the guide lists two kernel configurations:
- 256 bits at 450 MHz
- 512 bits at 225 MHz
The same principle can be applied to the V80 shell. Our draft implementation is available in commit 23d39e3f. We updated the linker to preserve the kernel’s AXI data width and generate a SmartConnect that performs asynchronous CDC and 512/256-bit width conversion. The static shell remains unchanged. No shell rebuild is required; only the application vbin needs to be regenerated. The kernel now uses a 512-bit interface targeting 200 MHz:
VRT Version: 1.0.0
Launching 64 perf kernels (perf_0 to perf_63)
Per-kernel buffer size: 512 MiB
Aggregate buffer footprint: 32.00 GiB
[2026-09-30 14:24:03.422] [INFO ] vrt::impl::Device::Device(const std::string&, const std::string&, bool, vrt::ProgramType): Programmed user clock to 197777777 Hz (target 197784810 Hz)
Write phase: launching 64 kernel(s)
Write phase time: 52 ms (654.08 GB/s aggregate)
Read phase: launching 64 kernel(s)
Read phase time: 48 ms (706.13 GB/s aggregate)
Combined read+write throughput: 681.40 GiB/s aggregate
Test passed
Scaling observations
We tested the modified 05_perf on V80 with 1–64 HBM channels, doubling the channel count between runs. With only a few channels, throughput closely matches the expected interface bandwidth. Aggregate throughput continues to increase with channel count, but efficiency decreases, particularly for writes. At 64 channels, read throughput reaches 95.8% of the interface ceiling, while write throughput reaches 88.7%.
We have not yet investigated the NoC design in depth, so the cause of this scaling loss remains unclear. Could it involve bandwidth constraints or contention within the NoC/HBM subsystem, memory-controller behavior, or other effects?
The solution is an RM-side workaround that leaves the static shell unchanged. If this feature would be useful upstream, we can open a pull request and help adapt the implementation for integration.
- Lenguaje dominante
- Tcl
- Estrellas
- 48
- Forks
- 23
- Merge medio
- 1 d 22 h
- PR fusionados (30 d)
- 7
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de Xilinx/SLASH
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 74/100
-
Dificultad 4/5 3-5 días Aptitud para principiantes 35/100
-
device bring up issue (v80-smi write-static-shell)Posiblemente ocupada @quetric la tomó hace 11 días. Abierto
Dificultad 4/5 3-5 días Aptitud para principiantes 35/100
-
The link is up and RX sees the packet, but the received packet is not counted as a good packet.Abierto
Dificultad 4/5 3-5 días Aptitud para principiantes 38/100
-
Dificultad 4/5 3-5 días Aptitud para principiantes 35/100
Todos los issues de Xilinx/SLASH
Issues similares
-
Dificultad 2/5 Menos de una hora Aptitud para principiantes 72/100
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 75/100
element-hq/lk-jwt-service#248 ·
Los mantenedores suelen responder en 1 día
-
area/install-update comp/cli duplicate P2 python:uv sweeper:risk-compatibility type/bug
Dificultad 1/5 Menos de una hora Aptitud para principiantes 62/100
NousResearch/hermes-agent#135440 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
pytorch/tensordict#1927 ·
Los mantenedores suelen responder en 1 día