vllm-project/vllm

[Feature]: Implement Concurrent Partial Prefills In V1 Engine

开放

#14,003 创建于 2025年2月28日

 (17 条评论) (0 个反应) (0 位负责人)Python (16,816 个派生)batch import
feature requesthelp wantedunstale

仓库指标

星标
 (80,034 个星标)
PR 合并指标
 (平均合并 3天 17小时) (30 天内合并 993 个 PR)

描述

🚀 The feature, motivation and pitch

In V0, we support concurrent partial prefills to avoid TTFT latency with long requests. Implement it in V1

cc @WoosukKwon

Alternatives

No response

Additional context

No response

Before submitting a new issue...

  • Make sure you already searched for relevant issues, and asked the chatbot living at the bottom right corner of the documentation page, which can answer lots of frequently asked questions.

贡献者指南