SIGSEGV in continuous batching when a streaming client disconnects mid-generation (block_manager.hpp:633 assertion, GPU, 2026.2.1)
Maintainer thường phản hồi trong vòng 2 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 45/100
Hướng nghiên cứu
Bắt đầu với openvino.genai/src/cpp/src/continuous_batching/cache/block_manager.hpp quanh dòng 633 và llm_executor.hpp quanh dòng 132, theo dõi việc hủy sau khi một client streaming ngắt kết nối. Tái hiện với yêu cầu streaming được cung cấp và thao tác đóng cưỡng bức; được xem là hoàn tất khi tiến trình vẫn hoạt động và các quá trình sinh khác đang chạy tiếp tục khi một client ngắt kết nối giữa chừng trong quá trình sinh.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Summary
A client that hard-closes its HTTP connection during a streaming
/v3/chat/completions generation reliably crashes the whole OVMS process within
seconds on Intel Arc GPUs. All other in-flight generations in the batch die with
it. Reproduced on demand; also observed 43× in 24 h under an agent fleet whose
callers time out and disconnect.
Environment
- Image:
docker.io/openvino/model_server:2026.2.1-gpu(stock, unmodified) - GPU: Intel Arc Pro B60 (xe driver) and Arc B580 — both affected, VM passthrough
- Model: Qwen3.6-27B INT4 OpenVINO IR (also seen with a 9B on the second card)
- Serving args:
--task=text_generation --target_device=GPU --plugin_config='{"ENABLE_CPU_PINNING":false}' --tool_parser qwen3coder --enable_tool_guided_generation true --reasoning_parser qwen3 --kv_cache_precision u8 --cache_size 4 --enable_prefix_caching true --cache_interval_multiplier 64 --max_num_seqs 16 --max_num_batched_tokens 4096
Reproducer (deterministic)
- POST a streaming chat completion (
"stream": true, long generation). - Read a few SSE chunks.
- Hard-close the client socket mid-stream (
socket.close(), no graceful shutdown). - Within ~90 s the server logs the assertion below and the process exits.
Observed on the very first attempt; kernel gained exactly one segfault entry.
Logs
[llm_executor][error][llm_executor.hpp:132] Error occurred in LLM executor:
Check 'm_block_table.count(seq_id) > 0' failed at
../../../../../repos/openvino.genai/src/cpp/src/continuous_batching/cache/block_manager.hpp:633
Kernel (55+ occurrences over days, all error 4 read faults):
ovms[…]: segfault at … ip …6f3 … error 4 in ovms
libstdc++.so.6.0.33[…] from mediapipe/NNNNN threads
Symbolized against the shipped binary (BuildID a1b35868566ed20139449c9674ff41b1104ec73c):
the faulting IP lands in spdlog::logger::sink_it_+0x43 — i.e. the process dies
while logging, consistent with the executor thread's error path racing the
request-drop path that has already freed the sequence's block table entry.
Impact
--max_num_seqs 16means one disconnected client kills up to 15 innocent
concurrent generations (All requests: 3; Scheduled requests: 3logged at death).- On a GPU deployment the recompile/reload takes 1–6 minutes per crash.
- Any latency spike becomes self-amplifying: slow generations → client timeouts →
disconnects → crash → colder cache → slower generations.
Notes
- The xe "CAT error → engine reset" events sometimes seen after the crash FOLLOW
the segfault by 300–460 ms and are absent for many crashes — downstream cleanup,
not the cause. - Not OOM (
memory.events oom_kill 0), not concurrency-proportional (a second
card at 2.3× the concurrent load crashes 51× less per stream-second — the
discriminating variable is the mid-stream disconnect rate). - 2026.1 release notes mention cancellation of requests from disconnected clients
(NPU); this looks like the GPU continuous-batching cancellation path freeing
seq_idfromm_block_tablewhile the executor still steps the sequence.
- Ngôn ngữ chính
- C++
- Star
- 932
- Fork
- 278
- Merge trung bình
- 2 ngày 15 giờ
- Pull request đã merge (30 ngày)
- 60
Chuẩn bị môi trường
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của openvinotoolkit/model_server
-
bug
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 68/100
openvinotoolkit/model_server#4609 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
-
`/v3/models` lists every model twice when `group_name` and `--idle_unload_timeout_seconds` are combinedCó thể đã có người làm @atobiszei đã nhận 3 ngày trước. Đang mở
openvinotoolkit/model_server#4604 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Idle unload never happens again if the client disconnects while a sleeping graph is waking upCó thể đã có người làm @atobiszei đã nhận 3 ngày trước. Đang mở
openvinotoolkit/model_server#4603 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 2 ngày
-
bug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 45/100
openvinotoolkit/model_server#4599 · 4 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 55/100
openvinotoolkit/model_server#4586 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
Tất cả issue của openvinotoolkit/model_server
Issue tương tự
-
WaterHeaterManagement: tank_percent feature reports wrong feature id (FeatureMap corruption)Đang mở
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
espressif/esp-matter#1867 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
Maintainer thường phản hồi trong vòng 1 ngày
-
cuAmpcorParameter.cpp (CPU pycuampcor) does not compile with libc++: EXIT_FAILURE is undeclaredĐang mở
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
isce-framework/isce3#387 ·
-
[request] opencv/5.0.0Đang mởupstream update
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
conan-io/conan-center-index#31055 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
Maintainer thường phản hồi trong vòng 1 ngày