Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

SIGSEGV in continuous batching when a streaming client disconnects mid-generation (block_manager.hpp:633 assertion, GPU, 2026.2.1)

未关闭
#4,428 21 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
45/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
活跃
技术栈
cpp
领域
ai, backend

调研方向

从 openvino.genai/src/cpp/src/continuous_batching/cache/block_manager.hpp 第 633 行附近和 llm_executor.hpp 第 132 行附近开始,跟踪 streaming 客户端断开连接后的取消流程。使用提供的 streaming 请求并执行强制关闭来复现;当客户端在生成过程中途断开连接时,进程仍保持运行且其他正在进行的生成继续运行,即表示完成。

由索引模型根据 Issue 内容生成。

描述

Summary

A client that hard-closes its HTTP connection during a streaming
/v3/chat/completions generation reliably crashes the whole OVMS process within
seconds on Intel Arc GPUs. All other in-flight generations in the batch die with
it. Reproduced on demand; also observed 43× in 24 h under an agent fleet whose
callers time out and disconnect.

Environment

  • Image: docker.io/openvino/model_server:2026.2.1-gpu (stock, unmodified)
  • GPU: Intel Arc Pro B60 (xe driver) and Arc B580 — both affected, VM passthrough
  • Model: Qwen3.6-27B INT4 OpenVINO IR (also seen with a 9B on the second card)
  • Serving args:
    --task=text_generation --target_device=GPU --plugin_config='{"ENABLE_CPU_PINNING":false}' --tool_parser qwen3coder --enable_tool_guided_generation true --reasoning_parser qwen3 --kv_cache_precision u8 --cache_size 4 --enable_prefix_caching true --cache_interval_multiplier 64 --max_num_seqs 16 --max_num_batched_tokens 4096

Reproducer (deterministic)

  1. POST a streaming chat completion ("stream": true, long generation).
  2. Read a few SSE chunks.
  3. Hard-close the client socket mid-stream (socket.close(), no graceful shutdown).
  4. Within ~90 s the server logs the assertion below and the process exits.

Observed on the very first attempt; kernel gained exactly one segfault entry.

Logs

[llm_executor][error][llm_executor.hpp:132] Error occurred in LLM executor:
Check 'm_block_table.count(seq_id) > 0' failed at
../../../../../repos/openvino.genai/src/cpp/src/continuous_batching/cache/block_manager.hpp:633

Kernel (55+ occurrences over days, all error 4 read faults):

ovms[…]: segfault at … ip …6f3 … error 4 in ovms
libstdc++.so.6.0.33[…] from mediapipe/NNNNN threads

Symbolized against the shipped binary (BuildID a1b35868566ed20139449c9674ff41b1104ec73c):
the faulting IP lands in spdlog::logger::sink_it_+0x43 — i.e. the process dies
while logging, consistent with the executor thread's error path racing the
request-drop path that has already freed the sequence's block table entry.

Impact

  • --max_num_seqs 16 means one disconnected client kills up to 15 innocent
    concurrent generations (All requests: 3; Scheduled requests: 3 logged at death).
  • On a GPU deployment the recompile/reload takes 1–6 minutes per crash.
  • Any latency spike becomes self-amplifying: slow generations → client timeouts →
    disconnects → crash → colder cache → slower generations.

Notes

  • The xe "CAT error → engine reset" events sometimes seen after the crash FOLLOW
    the segfault by 300–460 ms and are absent for many crashes — downstream cleanup,
    not the cause.
  • Not OOM (memory.events oom_kill 0), not concurrency-proportional (a second
    card at 2.3× the concurrent load crashes 51× less per stream-second — the
    discriminating variable is the mid-stream disconnect rate).
  • 2026.1 release notes mention cancellation of requests from disconnected clients
    (NPU); this looks like the GPU continuous-batching cancellation path freeing
    seq_id from m_block_table while the executor still steps the sequence.
主要语言
C++
星标
940
派生
278
平均合并
3 天 4 小时
30 天内合并 PR
68

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

openvinotoolkit/model_server 的其他 Issue

查看 openvinotoolkit/model_server 的全部 Issue

相似的 Issue

更多 C++ Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。