vllm-project/vllm-omni

[Good First Issue]: Add vendor-organized community recipes for “run model X on hardware Y for task Z” in vLLM-Omni

Open

#2.645 geöffnet am 9. Apr. 2026

Auf GitHub ansehen
 (59 Kommentare) (4 Reaktionen) (9 zugewiesene Personen)Python (1.067 Forks)github user discovery
good first issuehelp wantedhigh priority

Repository-Metriken

Stars
 (4.990 Stars)
PR-Merge-Metriken
 (PR-Metriken ausstehend)

Beschreibung

Community Recipes To-Do List

All supported hardware platforms (CUDA, ROCm, NPU, XPU) are welcome for every entry.

Legend: 🙋 help wanted | ⏳ PR raised | ✅ merged

Vendor Model Task Status Claimer PR Notes
Omni-Modality
Qwen Qwen3-Omni omni chat / serving 30B MoE (3B active)
Qwen Qwen2.5-Omni omni chat / speech @allgather #3151 7B / 3B
ByteDance BAGEL-7B-MoT image gen + understanding @GrayMiao123 MoT, ~42 GiB VRAM
ByteDance MammothModa2-Preview text-to-image 🙋 AR + DiT pipeline
inclusionAI Ming-flash-omni-2.0 omni-modal understanding @yuanheng-zhao #2890 MoE, 4-GPU TP
SNU AIDAS Dynin-Omni omni-modal (t2t/i2t/s2t/t2i/v2t/t2s) 🙋 3-stage token-based
SenseNova SenseNova-U1-8B-MoT omni-modal (T2I/I2I/I2T/T2T) @princepride #3319 MoE, ~36 GiB; think mode
Text-to-Image / Image Gen
Qwen Qwen-Image text-to-image @Semmer2 #2729 1024x1024, ~60 GiB
Qwen Qwen-Image-2512 text-to-image 🙋 @AbelSara updated variant
Qwen Qwen-Image-Edit image editing @yixiaoer text-guided editing
Qwen Qwen-Image-Layered layered image gen 🙋
Zhipu AI (GLM) GLM-Image image gen / editing @friden-zhang 2-stage AR+DiT
Tongyi Z-Image-Turbo text-to-image (distilled) @MrDongsls #3283 4-9 steps, ~25 GiB; 2x RTX 5880 48GB
Stepfun NextStep-1.1 text-to-image @GrayMiao123 dual-level CFG, 512x512
Meituan LongCat-Image text-to-image 🙋 @Mashirona 1024x1024
Meituan LongCat-Image-Edit image editing 🙋
OvisAI Ovis-Image text-to-image 🙋 1024x1024
OmniGen2 OmniGen2 text-to-image 🙋 @Jerry2423 @Yuyi-Ao @jackywangno007-cyber ~20 GiB
Stability AI Stable-Diffusion-3.5 text-to-image 🙋 @yangyonggit 1024x1024
Black Forest Labs FLUX.1-dev text-to-image 🙋 ~78 GiB
Black Forest Labs FLUX.1-schnell text-to-image (fast) 🙋 fast distilled variant
Black Forest Labs FLUX.2-dev text-to-image 🙋 needs CPU offload on 80GiB
Black Forest Labs FLUX.2-klein-4B text-to-image 🙋 distilled, 4B params
Black Forest Labs FLUX.2-klein-9B text-to-image 🙋 distilled, 9B params
Tencent HunyuanImage-3.0 image gen / understanding @Bounty-hunter #2495 MoE, dual task
Baidu ERNIE-Image text-to-image @RuixiangMa #2861 1024x1024, ~23 GiB
Text-to-Video / Image-to-Video
Tencent-Hunyuan HunyuanVideo-1.5-T2V text-to-video @allgather #3152 480p path on 1xA100; needs FP8 + VAE tiling for 720p
Tencent-Hunyuan HunyuanVideo-1.5-I2V image-to-video @xldeng-chn 1xA100 80GB
Wan-AI Wan2.2-T2V-A14B text-to-video @lengrongfu #3018 720x1280, 81 frames, ~60 GiB
Wan-AI Wan2.2-TI2V-5B text+image-to-video 🙋 unified, 5B params
Wan-AI Wan2.2-I2V-A14B image-to-video MoE
Wan-AI Wan2.1-VACE video creation (T2V/I2V/FLF2V) 🙋 1.3B / 14B variants
Wan-AI Wan2.2-S2V-14B speech-to-video (image+audio→video) @xuechendi #2751 720p, 81 frames; TP=2
Lightricks LTX-2-T2V text-to-video @fywc #3294 512x768, 121 frames
Lightricks LTX-2-I2V image-to-video @fywc #3294 same model as T2V
Lightricks LTX-2.3 text-to-video + audio @oglok #2893 22B, ~62 GiB peak; BWE 48kHz vocoder
Helios Helios-Base T2V / I2V / V2V @lengrongfu stage 1 only; 8xH800 or 8xH200
Helios Helios-Mid video denoising 🙋 @lengrongfu multi-stage pyramid; 8xH800 or 8xH200
Helios Helios-Distilled video (few-step) 🙋 @lengrongfu DMD distillation; 8xH800 or 8xH200
SII-GAIR MagiHuman video+audio with lip sync 🙋 TP=4 for 80GB, DiT MoE+T5
XuGuo699 DreamID-Omni I2V with identity preservation 🙋 @fywc ~72 GiB
Text-to-Speech / Audio
Qwen Qwen3-TTS TTS serving @chzhang2021 #3130 12Hz variants (CustomVoice, VoiceDesign, Base)
FunAudioLLM CosyVoice3 TTS @Moore-Z #3486 2-stage: talker + flow-matching, 0.5B
FishAudio Fish Speech S2 Pro TTS / voice cloning @menjiantong #3193 4B dual-AR, 44.1kHz
Mistral AI Voxtral TTS TTS with voice presets 🙋 4B, voice cloning
K2-FSA OmniVoice multilingual TTS (600+ lang) 🙋 @fray1024 zero-shot, Qwen3-0.6B backbone; 1xL20 48GB
OpenBMB VoxCPM2 TTS (diffusion AR) 🙋 @wjinxu 2B, 48kHz, 30+ lang
Xiaomi (MiMo) MiMo-Audio audio gen / TTS / ASR / dialogue 🙋 multi-task audio 7B
Stability AI Stable-Audio-Open audio gen (music, SFX) 🙋 @Ronnie-Rui diffusion-based
HKUST AudioX T2A / V2A / TV2A / T2M / V2M / TV2M @zhangj1an #2077 diffusion maf-mmdit; ~97 GiB VRAM
Tencent Covo-Audio-Chat audio chat (audio in, text+audio out) @Dnoob #2293 7B LLM + BigVGAN; ~20 GiB

How to contribute (step by step)

  1. Claim a row — comment on this issue with the row + your hardware (e.g. Qwen3-TTS, 1xL40S 48GB). Maintainer will mark the row ⏳ and assign you.
  2. Copy the templatecp recipes/_template.md recipes/<Vendor>/<Model>.md (template merged in #2646).
  3. Fill the required sections — model + task, tested HW, env (CUDA/ROCm/CANN versions), launch command, verification command + expected output, important flags, known limitations, links back to docs/ and examples/.
  4. Test end-to-end on the hardware you claimed — recipes that aren't personally validated should not be merged.
  5. Open PR titled [Recipe] <Vendor>/<Model> (or [Recipe] <Vendor>/<Model> (<HW>) if adding a HW section to an existing recipe), link this issue, and paste the launch command + verification output in the PR body. Include a test plan describing what you verified (e.g. specific prompts, expected outputs, latency/throughput numbers) and the test results showing the actual output or logs.

The initial recipes/ template has been merged #2646. We now welcome community contributions for practical "run model X on hardware Y for task Z" recipes in vllm-omni.

Please use the merged template as the starting point and organize recipes by model vendor, for example:

recipes/
  Qwen/
    Qwen3-Omni.md
    Qwen3-TTS.md
    Qwen2.5-Omni.md
  Tencent-Hunyuan/
    HunyuanVideo.md
  GLM/
    GLM-Image.md
  MiMo/
    MiMo-Audio.md
  Wan-AI/
    Wan2.2-T2V.md
  ...

Each recipe should include:

  • Model name and task
  • Tested hardware configuration
  • Required environment
  • Launch command
  • Verification command or expected output
  • Important flags or stage configs
  • Known limitations
  • Links back to canonical docs/ and runnable examples/

High Priority Starter Recipes

  • recipes/Qwen/Qwen3-TTS.md

    • Suggested scope: text-to-speech serving with Qwen3-TTS
    • Suggested hardware sections: CUDA GPU, NPU if available
    • Link to existing Qwen3-TTS docs/examples where possible
  • recipes/Qwen/Qwen2.5-Omni.md

    • Suggested scope: omni-modal chat or speech interaction
    • Suggested hardware sections: CUDA GPU, NPU if available
    • Link to existing Qwen2.5-Omni docs/examples where possible
  • recipes/Tencent-Hunyuan/HunyuanVideo.md

    • Suggested scope: text-to-video or image-to-video generation
    • Include tested VRAM requirements and performance notes where possible
  • recipes/GLM/GLM-Image.md

    • Suggested scope: image generation with GLM-Image
    • Include recommended generation settings and validation output
  • recipes/MiMo/MiMo-Audio.md

    • Suggested scope: audio generation or speech-related serving
    • Include model-specific setup and verification notes

Hardware Coverage We Want

  • NVIDIA CUDA recipes

    • Example targets: A100 80GB, H100, L40S, RTX 4090 where applicable
  • AMD ROCm recipes

    • Include ROCm version, GPU model, and any known caveats
  • Intel XPU recipes

    • Suggested devices from community feedback:
    • Intel Arc Pro B50, 16GB
    • Intel Arc Pro B60, 24GB
    • Intel Arc Pro B70, 32GB
  • Huawei NPU recipes

    • Include CANN/runtime versions and model-specific limitations

Good First Contributions

  • Add a recipe for a model you have successfully run locally
  • Add another hardware section to an existing recipe
  • Add verification output to an existing recipe
  • Link an existing recipe to the relevant docs/ and examples/
  • Improve clarity around memory usage, flags, or known limitations

Contribution Guidelines

When opening a recipe PR, please include:

  • The exact command you used
  • Hardware details
  • Software/runtime versions
  • Whether the recipe was personally tested
  • Test plan: what you verified (specific prompts, expected outputs, latency/throughput numbers)
  • Test results: actual output or logs from your run
  • Any limitations or assumptions
  • Links to relevant examples or docs

Recipes should not replace canonical documentation. They should act as practical, community-maintained runbooks that point users to the right docs and runnable examples.


Motivation.

We'd like to propose a new community-maintained recipes/ area in vllm-project/vllm-omni to answer a recurring user question:

How do I run model X on hardware Y for task Z?

Today, users often struggle to find the best path for a concrete deployment scenario in vLLM-Omni.

There are a few reasons:

  1. We currently do not have a dedicated place for operational runbooks that map a specific model, hardware target, and task to a known-good setup.
  2. The current examples/ layout mixes model-specific and task-specific directories, which can be confusing for discovery.
  3. We want to align the user experience with vllm-project/recipes, which already provides this style of practical guide for upstream vLLM.

Proposed Change.

Introduce a top-level recipes/ directory in vLLM-Omni for community-maintained runbooks.

To align with upstream vllm-project/recipes, recipes should be grouped by model vendor at the top level, with one Markdown file per model family by default.

Example direction:

recipes/
  Qwen/
    Qwen3-Omni.md
    Qwen3-TTS.md
    Qwen2.5-Omni.md
  Tencent-Hunyuan/
    HunyuanVideo.md
  GLM/
    GLM-Image.md
  MiMo/
    MiMo-Audio.md

Within each model doc, we should include multiple hardware-specific sections in the same Markdown file, following the structure used by the upstream DeepSeek guide: https://github.com/vllm-project/recipes/blob/main/DeepSeek/DeepSeek-V3.md

For example, a single recipe doc could contain sections such as:

  • 1x A100 80GB
  • 2x L40S
  • 4x H100
  • ROCm / NPU variants where applicable

Each section would describe:

  • supported task(s)
  • tested hardware
  • required environment
  • launch commands
  • important flags / stage configs
  • verification steps
  • known limitations

Design Principles

  • Align the top-level organization with vllm-project/recipes where practical, so discovery feels familiar across vLLM and vLLM-Omni.
  • Keep one Markdown file per model family by default.
  • Keep multiple hardware configurations inside the same recipe document unless the document becomes too large or hard to maintain.
  • Use recipes for practical "known-good setup" guidance, not as the canonical source of product documentation.

Relationship to examples/

The current examples/ folder is still valuable, but it serves a different purpose:

  • examples/: runnable code and scripts
  • docs/: canonical documentation
  • recipes/: practical "known-good setup" guides for concrete user scenarios

One benefit of adding recipes/ is that it may reduce pressure to make examples/ itself carry all discovery and onboarding needs.

This also gives us a cleaner answer to users who ask for a concrete deployment path without requiring them to infer it from a mix of task-oriented and model-oriented example folders.

Scope

This proposal is for community recipes only.

It is not intended to replace:

  • canonical product documentation under docs/
  • source-of-truth runnable examples under examples/

Instead, recipes would act as a user-oriented entry point that links back to those canonical materials.

Open Questions

  1. Should recipe ownership be fully community-maintained, or should each recipe have one or two named maintainers?
  2. Should recipes live only in the repo at first, or also be surfaced in the documentation site under a Community section?
  3. Should we add a lightweight recipe template from the beginning to keep structure consistent?
  4. Should any future cleanup of examples/ be handled in a separate RFC after recipes/ is established?

Initial Success Criteria

  • A new user can quickly find a concrete recipe for a target model, hardware, and task.
  • Recipes follow a consistent structure.
  • Recipes are organized by vendor in a way that feels familiar to users of upstream vllm-project/recipes.
  • Recipes link to examples/ and docs/ instead of duplicating canonical content.
  • The approach improves user experience without adding confusion about where canonical documentation lives.

Suggested First Recipes

  • recipes/Qwen/Qwen3-Omni.md
  • recipes/Qwen/Qwen3-TTS.md
  • recipes/Qwen/Qwen2.5-Omni.md
  • recipes/Tencent-Hunyuan/HunyuanVideo.md
  • recipes/GLM/GLM-Image.md
  • recipes/MiMo/MiMo-Audio.md

Feedback welcome on structure, ownership, and how tightly we should align with upstream vllm-project/recipes.

Feedback Period.

CC List.

vllm-omni maintainer team @ywang96 @Gaohan123 @ZJY0516 @princepride @lishunyang12 .....

Any Other Things.

Contributor Guide