SIGSEGV in continuous batching when a streaming client disconnects mid-generation (block_manager.hpp:633 assertion, GPU, 2026.2.1)
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 45/100
Línea de trabajo
Comienza con openvino.genai/src/cpp/src/continuous_batching/cache/block_manager.hpp alrededor de la línea 633 y llm_executor.hpp alrededor de la línea 132, siguiendo la cancelación después de que un cliente de streaming se desconecte. Reproduce el problema con la solicitud de streaming proporcionada y un cierre forzado; se considera completado cuando el proceso permanece activo y las demás generaciones en curso continúan cuando un cliente se desconecta a mitad de la generación.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Summary
A client that hard-closes its HTTP connection during a streaming
/v3/chat/completions generation reliably crashes the whole OVMS process within
seconds on Intel Arc GPUs. All other in-flight generations in the batch die with
it. Reproduced on demand; also observed 43× in 24 h under an agent fleet whose
callers time out and disconnect.
Environment
- Image:
docker.io/openvino/model_server:2026.2.1-gpu(stock, unmodified) - GPU: Intel Arc Pro B60 (xe driver) and Arc B580 — both affected, VM passthrough
- Model: Qwen3.6-27B INT4 OpenVINO IR (also seen with a 9B on the second card)
- Serving args:
--task=text_generation --target_device=GPU --plugin_config='{"ENABLE_CPU_PINNING":false}' --tool_parser qwen3coder --enable_tool_guided_generation true --reasoning_parser qwen3 --kv_cache_precision u8 --cache_size 4 --enable_prefix_caching true --cache_interval_multiplier 64 --max_num_seqs 16 --max_num_batched_tokens 4096
Reproducer (deterministic)
- POST a streaming chat completion (
"stream": true, long generation). - Read a few SSE chunks.
- Hard-close the client socket mid-stream (
socket.close(), no graceful shutdown). - Within ~90 s the server logs the assertion below and the process exits.
Observed on the very first attempt; kernel gained exactly one segfault entry.
Logs
[llm_executor][error][llm_executor.hpp:132] Error occurred in LLM executor:
Check 'm_block_table.count(seq_id) > 0' failed at
../../../../../repos/openvino.genai/src/cpp/src/continuous_batching/cache/block_manager.hpp:633
Kernel (55+ occurrences over days, all error 4 read faults):
ovms[…]: segfault at … ip …6f3 … error 4 in ovms
libstdc++.so.6.0.33[…] from mediapipe/NNNNN threads
Symbolized against the shipped binary (BuildID a1b35868566ed20139449c9674ff41b1104ec73c):
the faulting IP lands in spdlog::logger::sink_it_+0x43 — i.e. the process dies
while logging, consistent with the executor thread's error path racing the
request-drop path that has already freed the sequence's block table entry.
Impact
--max_num_seqs 16means one disconnected client kills up to 15 innocent
concurrent generations (All requests: 3; Scheduled requests: 3logged at death).- On a GPU deployment the recompile/reload takes 1–6 minutes per crash.
- Any latency spike becomes self-amplifying: slow generations → client timeouts →
disconnects → crash → colder cache → slower generations.
Notes
- The xe "CAT error → engine reset" events sometimes seen after the crash FOLLOW
the segfault by 300–460 ms and are absent for many crashes — downstream cleanup,
not the cause. - Not OOM (
memory.events oom_kill 0), not concurrency-proportional (a second
card at 2.3× the concurrent load crashes 51× less per stream-second — the
discriminating variable is the mid-stream disconnect rate). - 2026.1 release notes mention cancellation of requests from disconnected clients
(NPU); this looks like the GPU continuous-batching cancellation path freeing
seq_idfromm_block_tablewhile the executor still steps the sequence.
- Lenguaje dominante
- C++
- Estrellas
- 932
- Forks
- 278
- Merge medio
- 2 d 23 h
- PR fusionados (30 d)
- 68
Preparar el entorno
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de openvinotoolkit/model_server
-
Dificultad 3/5 1-2 días Aptitud para principiantes 58/100
openvinotoolkit/model_server#4613 ·
Los mantenedores suelen responder en 1 día
-
enhancement
Dificultad 3/5 1-2 días Aptitud para principiantes 68/100
openvinotoolkit/model_server#4609 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
`/v3/models` lists every model twice when `group_name` and `--idle_unload_timeout_seconds` are combinedPosiblemente ocupada @atobiszei la tomó hace 4 días. Abierto
openvinotoolkit/model_server#4604 · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
Idle unload never happens again if the client disconnects while a sleeping graph is waking upPosiblemente ocupada @atobiszei la tomó hace 4 días. Abierto
openvinotoolkit/model_server#4603 · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 4/5 3-5 días Aptitud para principiantes 45/100
openvinotoolkit/model_server#4599 · 4 comentarios ·
Los mantenedores suelen responder en 1 día
Todos los issues de openvinotoolkit/model_server
Issues similares
-
`enzymexla.linalg.lu` lowering fails for a tall matrix: the permutation is built with the pivot typeAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
EnzymeAD/Enzyme-JAX#3286 ·
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
apache/iceberg-cpp#973 ·
Los mantenedores suelen responder en 1 día
-
Add c++23 mapping to nvccAbiertofeature request
Dificultad 1/5 Menos de una hora Aptitud para principiantes 86/100
Los mantenedores suelen responder en 2 días
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
Los mantenedores suelen responder en 1 día