vllm-project/vllm
[Feature]: Implement Concurrent Partial Prefills In V1 Engine
オープン
#14,003 opened on 2025/02/28
feature requesthelp wantedunstale
Repository metrics
- Stars
- (80,034 個のスター)
- PR merge metrics
- (平均マージ 3d 17h) (30d で 993 merged PRs)
説明
🚀 The feature, motivation and pitch
In V0, we support concurrent partial prefills to avoid TTFT latency with long requests. Implement it in V1
cc @WoosukKwon
Alternatives
No response
Additional context
No response
Before submitting a new issue...
- Make sure you already searched for relevant issues, and asked the chatbot living at the bottom right corner of the documentation page, which can answer lots of frequently asked questions.