flashinfer-ai/flashinfer

flashinfer.nvfp4_quantize: make CuTe-DSL backend's `a_global_sf` arg a host side arg (to align CuTe-DSL backend with CUDA backend)

クローズ

#4,112 opened on 2026/07/23

 (2 件のコメント) (0 件のリアクション) (1 人の担当者)Python (1,031 件のフォーク)github user discovery
good first issueop: misc

Repository metrics

Stars
 (5,756 個のスター)
PR merge metrics
 (平均マージ 11d 23h) (30d で 186 merged PRs)

説明

As reported: "In flashinfer.nvfp4_quantize, the type of the argument a_global_sf is inconsistent. For cuda, it is a host side value; for cute-dsl, it is a device side tensor. The function description is not clear enough. We should make it host side and passed as a function argument.

What we might be able to do relatively easily is to change the cute-dsl backend kernel to accept either (a) a device side tensor of size 1 and dtype float32 (what we support today); or (b) a single float value that trivially comes from host.  That way it would be backwards compatible for those who might be using the device-side tensor."

コントリビューターガイド