关于实时模式的vllm解码并发问题
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 35/100
- Issue type
- Refactor
- Clarity
- Needs clarification
- Activity status
- Active
- Tech stack
- python
- Domain
- backend, performance
Research direction
Start in the realtime_ws flow around RealtimeBatchingEngine, handle_client, and run_session_work; trace the asyncio.to_thread call into synchronous LLM.generate. Run the existing concurrency benchmark, which currently covers up to 8 clients, and compare it with higher session counts. Done means a benchmark-backed decision and, if pursued, verified per-session partial ordering and latency under the target concurrency.
Written by the indexing model from the issue text.
Description
【realtime_ws】并发 streaming session 的vllm解码
测试中并发 streaming session 增多时,首条结果延迟明显上升,超过一定并发后结果趋于停滞。
在约 15+ 个并发 session 下,首字延迟 p50 从 10 个 session 时的 ~1.5s 涨到 ~2.2s(p95
到 ~66s),并随并发继续恶化。以上数值仅供参考,不代表精确基准。
似乎RealtimeBatchingEngine把并发 session 的请求统一汇到单个 worker 线程,底层调用同步 `LLM.generate。实时流式里每个 session 的 handle_client循环要等自己的 partial 返回之后才发下一帧(run_session_work → asyncio.to_thread →阻塞 generate),语音请求在vad处理后请求天然稀疏,batch_wait_ms=10ms的攒批窗口几乎凑不齐。于是请求稀疏 → 批凑不满 → 单条串行 → 客户端等更久 → 更稀疏,形成自锁,吞吐受限。这是我的推测解释,未必准确,供参考。
目前benchmark 文档只到 8 clients,若未来有意支持更高并发,是否可以考虑把解码器切到 AsyncLLMEngine(vLLM v1 的vllm.v1.engine.async_llm.AsyncLLM,同样 enable_prompt_embeds=True),每个 session 的partial / 锁句作独立 request_id提交:引擎按 step 在卡上合批、同时把每条结果返回给各自调用方,天然保住每个 session 的顺序与 partial 中间态,吞吐也随之恢复。是否值得,请各位大佬按定位评估;我这边物理条件有限,测试结果未必准确,也很难支持更高并发下的吞吐测试,若感兴趣可在高并发下尝试验证。
- Dominant language
- Python
- Stars
- 20.5k
- Forks
- 2.1k
- Avg merge
- 15h 21m
- Merged PRs (30d)
- 125
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from modelscope/FunASR
-
Difficulty 1/5 Under an hour Newbie friendliness 88/100
modelscope/FunASR#3730 · 2 comments ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
modelscope/FunASR#3704 · 1 comment ·
Maintainers usually reply within 1 day
-
bug needs feedback
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
modelscope/FunASR#3401 · 2 comments ·
Maintainers usually reply within 1 day
-
needs feedback
Difficulty 4/5 3-5 days Newbie friendliness 45/100
modelscope/FunASR#3739 · 1 comment ·
Maintainers usually reply within 1 day
-
bug needs triage
Difficulty 4/5 3-5 days Newbie friendliness 30/100
modelscope/FunASR#3727 · 2 comments ·
Maintainers usually reply within 1 day
All issues in modelscope/FunASR
Similar issues
-
repo-audit
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
scverse/repo-health#20 ·
Maintainers usually reply within 1 day
-
/context/prime scope override double-prefixes an entity-ref project and drops its scoped memoriesOpen
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
phasespace-labs/palinode#232 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
collective/icalendar#1858 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
langflow-ai/langflow#15496 ·
Maintainers usually reply within 1 day