SIGSEGV in continuous batching when a streaming client disconnects mid-generation (block_manager.hpp:633 assertion, GPU, 2026.2.1)
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 45/100
Direzione di ricerca
Inizia da openvino.genai/src/cpp/src/continuous_batching/cache/block_manager.hpp intorno alla riga 633 e llm_executor.hpp intorno alla riga 132, seguendo la cancellazione dopo la disconnessione di un client di streaming. Riproduci il problema con la richiesta di streaming fornita e una chiusura forzata; il lavoro è completato quando il processo rimane attivo e le altre generazioni in corso continuano quando un client si disconnette durante la generazione.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Summary
A client that hard-closes its HTTP connection during a streaming
/v3/chat/completions generation reliably crashes the whole OVMS process within
seconds on Intel Arc GPUs. All other in-flight generations in the batch die with
it. Reproduced on demand; also observed 43× in 24 h under an agent fleet whose
callers time out and disconnect.
Environment
- Image:
docker.io/openvino/model_server:2026.2.1-gpu(stock, unmodified) - GPU: Intel Arc Pro B60 (xe driver) and Arc B580 — both affected, VM passthrough
- Model: Qwen3.6-27B INT4 OpenVINO IR (also seen with a 9B on the second card)
- Serving args:
--task=text_generation --target_device=GPU --plugin_config='{"ENABLE_CPU_PINNING":false}' --tool_parser qwen3coder --enable_tool_guided_generation true --reasoning_parser qwen3 --kv_cache_precision u8 --cache_size 4 --enable_prefix_caching true --cache_interval_multiplier 64 --max_num_seqs 16 --max_num_batched_tokens 4096
Reproducer (deterministic)
- POST a streaming chat completion (
"stream": true, long generation). - Read a few SSE chunks.
- Hard-close the client socket mid-stream (
socket.close(), no graceful shutdown). - Within ~90 s the server logs the assertion below and the process exits.
Observed on the very first attempt; kernel gained exactly one segfault entry.
Logs
[llm_executor][error][llm_executor.hpp:132] Error occurred in LLM executor:
Check 'm_block_table.count(seq_id) > 0' failed at
../../../../../repos/openvino.genai/src/cpp/src/continuous_batching/cache/block_manager.hpp:633
Kernel (55+ occurrences over days, all error 4 read faults):
ovms[…]: segfault at … ip …6f3 … error 4 in ovms
libstdc++.so.6.0.33[…] from mediapipe/NNNNN threads
Symbolized against the shipped binary (BuildID a1b35868566ed20139449c9674ff41b1104ec73c):
the faulting IP lands in spdlog::logger::sink_it_+0x43 — i.e. the process dies
while logging, consistent with the executor thread's error path racing the
request-drop path that has already freed the sequence's block table entry.
Impact
--max_num_seqs 16means one disconnected client kills up to 15 innocent
concurrent generations (All requests: 3; Scheduled requests: 3logged at death).- On a GPU deployment the recompile/reload takes 1–6 minutes per crash.
- Any latency spike becomes self-amplifying: slow generations → client timeouts →
disconnects → crash → colder cache → slower generations.
Notes
- The xe "CAT error → engine reset" events sometimes seen after the crash FOLLOW
the segfault by 300–460 ms and are absent for many crashes — downstream cleanup,
not the cause. - Not OOM (
memory.events oom_kill 0), not concurrency-proportional (a second
card at 2.3× the concurrent load crashes 51× less per stream-second — the
discriminating variable is the mid-stream disconnect rate). - 2026.1 release notes mention cancellation of requests from disconnected clients
(NPU); this looks like the GPU continuous-batching cancellation path freeing
seq_idfromm_block_tablewhile the executor still steps the sequence.
- Lingua principale
- C++
- Stelle
- 932
- Fork
- 278
- Merge medio
- 2g 23h
- PR unite (30g)
- 68
Preparare l'ambiente
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di openvinotoolkit/model_server
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 58/100
openvinotoolkit/model_server#4613 ·
I maintainer di solito rispondono entro 1 giorno
-
enhancement
Difficoltà 3/5 1-2 giorni Idoneità per principianti 68/100
openvinotoolkit/model_server#4609 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
`/v3/models` lists every model twice when `group_name` and `--idle_unload_timeout_seconds` are combinedForse già presa @atobiszei l’ha presa 4 giorni fa. Aperta
openvinotoolkit/model_server#4604 · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
Idle unload never happens again if the client disconnects while a sleeping graph is waking upForse già presa @atobiszei l’ha presa 4 giorni fa. Aperta
openvinotoolkit/model_server#4603 · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 4/5 3-5 giorni Idoneità per principianti 45/100
openvinotoolkit/model_server#4599 · 4 commenti ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di openvinotoolkit/model_server
Issue simili
-
bug
Difficoltà 1/5 1-3 ore Idoneità per principianti 88/100
isl-org/Open3D#7585 · 1 commento ·
I maintainer di solito rispondono entro 2 giorni
-
Unconfirmed bug
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
luanti-org/luanti#17605 · 1 commento ·
I maintainer di solito rispondono entro 2 giorni
-
area: config area: firmware priority: P2 - medium size: S type: bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
Mizithra/ActiveTerrain#16 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
grumpycoders/pcsx-redux#2171 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
I maintainer di solito rispondono entro 2 giorni