vllm-project/vllm-omni

[Aligned with vllm 0.24.0 now][RFC]: vLLM-Omni 2026 Q2 Roadmap

Geschlossen

#2.136 geöffnet am 24.03.2026

 (11 Kommentare) (7 Reaktionen) (1 zugewiesene Person)Python (1.067 Forks)github user discovery
good first issuehelp wantedroadmap

Repository-Metriken

Stars
 (4.990 Sterne)
PR-Merge-Metriken
 (PR-Metriken ausstehend)

Beschreibung

Motivation.

This issue outlines the development roadmap for vllm-omni in Q2 2026. We are collecting your feedbacks and finalize it later.

Proposed Change.

1) Model Support

  • Omni Models #2207
  • TTS Models #2115
  • Diffusion Models #2226 including world models and vla models

2) Feature Support

  • Quantization #1854
  • Prefix Cache #1184
  • KV offload #1330
  • Model Runner V2 #1770
  • Streaming audio/video input/output #2208
  • RL integration with veRL https://github.com/verl-project/verl/issues/5755
  • Diffusion Optimizations #1217
  • Diffusion continuous batching: #1769
  • Diffusers backend #2403

3) Large Scale Deployment (Disaggregated Serving)

  • EPDG disaggregation: x(E)y(P/D)zG #2336
  • targeting models: Bagel/HY-Image3.0/Qwen-Omni/Ming-Omni

4) Hardware Support

  • ROCm (AMD): #2413
  • Intel XPU: #2574
  • Ascend NPU: #2223
  • MUSA: #2347

5) CI/CD & Usability

  • Quality Gates: L1-L5 level tests enhancement&refactor #2299 #4197 #4254
  • Benchmarking performance: vllm-omni-kanban
  • Developer Usability: vllm-omni-skills
  • User guidance: vllm/recipes

Feedback Period.

No response

CC List.

@ywang96 @Gaohan123 @youkaichao @Isotr0py @ZJY0516 @david6666666 @SamitHuang @linyueqian @wtomin @tzhouam @gcanlin @tjtanaa @xuechendi

Any Other Things.

No response

Before submitting a new issue...

  • Make sure you already searched for relevant issues, and asked the chatbot living at the bottom right corner of the documentation page, which can answer lots of frequently asked questions.

Contributor Guide