vllm-project/vllm-omni

Consider using pre-built Flash Attention kernels via `kernels`

Aperta

#3922 aperta il 28 mag 2026

 (5 commenti) (0 reazioni) (0 assegnatari)Python (1067 fork)github user discovery
help wanted

Metriche repository

Star
 (4990 stelle)
Metriche merge PR
 (Metriche PR in attesa)

Descrizione

Hey,

I am Sayak from the Kernels team at Hugging Face. I noticed that this project uses Flash Attention which includes a long build time. We ship pre-built binaries (which provide bit-exact outputs as the upstream) and thereby, we make it easy to use.

Using FA3 on a supported machine is as easy as:

# make sure `kernels` is installed: `pip install -U kernels`
from kernels import get_kernel

kernel_module = get_kernel("kernels-community/flash-attn3")
flash_attn_func = kernel_module.flash_attn_func

flash_attn_func(...)

Let us know if you'd be interested in this and and we'd be happy to provide a draft of how it would look in your repo.

Guida contributor