Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

关于实时模式的vllm解码并发问题

Open
#3,732 1 comment 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
35/100
Issue type
Refactor
Clarity
Needs clarification
Activity status
Active
Tech stack
python

Research direction

Start in the realtime_ws flow around RealtimeBatchingEngine, handle_client, and run_session_work; trace the asyncio.to_thread call into synchronous LLM.generate. Run the existing concurrency benchmark, which currently covers up to 8 clients, and compare it with higher session counts. Done means a benchmark-backed decision and, if pursued, verified per-session partial ordering and latency under the target concurrency.

Written by the indexing model from the issue text.

Description

benchmark needs feedback question
【realtime_ws】并发 streaming session 的vllm解码

测试中并发 streaming session 增多时,首条结果延迟明显上升,超过一定并发后结果趋于停滞。
在约 15+ 个并发 session 下,首字延迟 p50 从 10 个 session 时的 ~1.5s 涨到 ~2.2s(p95
到 ~66s),并随并发继续恶化。以上数值仅供参考,不代表精确基准。

似乎RealtimeBatchingEngine把并发 session 的请求统一汇到单个 worker 线程,底层调用同步 `LLM.generate。实时流式里每个 session 的 handle_client循环要等自己的 partial 返回之后才发下一帧(run_session_work → asyncio.to_thread →阻塞 generate),语音请求在vad处理后请求天然稀疏,batch_wait_ms=10ms的攒批窗口几乎凑不齐。于是请求稀疏 → 批凑不满 → 单条串行 → 客户端等更久 → 更稀疏,形成自锁,吞吐受限。这是我的推测解释,未必准确,供参考。

目前benchmark 文档只到 8 clients,若未来有意支持更高并发,是否可以考虑把解码器切到 AsyncLLMEngine(vLLM v1 的vllm.v1.engine.async_llm.AsyncLLM,同样 enable_prompt_embeds=True),每个 session 的partial / 锁句作独立 request_id提交:引擎按 step 在卡上合批、同时把每条结果返回给各自调用方,天然保住每个 session 的顺序与 partial 中间态,吞吐也随之恢复。是否值得,请各位大佬按定位评估;我这边物理条件有限,测试结果未必准确,也很难支持更高并发下的吞吐测试,若感兴趣可在高并发下尝试验证。

Dominant language
Python
Stars
20.5k
Forks
2.1k
Avg merge
15h 21m
Merged PRs (30d)
125

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from modelscope/FunASR

All issues in modelscope/FunASR

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.