vllm-project/vllm-omni

Consider using pre-built Flash Attention kernels via `kernels`

オープン

#3,922 opened on 2026/05/28

 (5 件のコメント) (0 件のリアクション) (0 人の担当者)Python (1,067 件のフォーク)github user discovery
help wanted

Repository metrics

Stars
 (4,990 個のスター)
PR merge metrics
 (PR metrics pending)

説明

Hey,

I am Sayak from the Kernels team at Hugging Face. I noticed that this project uses Flash Attention which includes a long build time. We ship pre-built binaries (which provide bit-exact outputs as the upstream) and thereby, we make it easy to use.

Using FA3 on a supported machine is as easy as:

# make sure `kernels` is installed: `pip install -U kernels`
from kernels import get_kernel

kernel_module = get_kernel("kernels-community/flash-attn3")
flash_attn_func = kernel_module.flash_attn_func

flash_attn_func(...)

Let us know if you'd be interested in this and and we'd be happy to provide a draft of how it would look in your repo.

コントリビューターガイド