Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

SIGSEGV in continuous batching when a streaming client disconnects mid-generation (block_manager.hpp:633 assertion, GPU, 2026.2.1)

Aperta
#4,428 19 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
45/100
Tipo di issue
Bug
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
cpp
Ambito
ai, backend

Direzione di ricerca

Inizia da openvino.genai/src/cpp/src/continuous_batching/cache/block_manager.hpp intorno alla riga 633 e llm_executor.hpp intorno alla riga 132, seguendo la cancellazione dopo la disconnessione di un client di streaming. Riproduci il problema con la richiesta di streaming fornita e una chiusura forzata; il lavoro è completato quando il processo rimane attivo e le altre generazioni in corso continuano quando un client si disconnette durante la generazione.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Summary

A client that hard-closes its HTTP connection during a streaming
/v3/chat/completions generation reliably crashes the whole OVMS process within
seconds on Intel Arc GPUs. All other in-flight generations in the batch die with
it. Reproduced on demand; also observed 43× in 24 h under an agent fleet whose
callers time out and disconnect.

Environment

  • Image: docker.io/openvino/model_server:2026.2.1-gpu (stock, unmodified)
  • GPU: Intel Arc Pro B60 (xe driver) and Arc B580 — both affected, VM passthrough
  • Model: Qwen3.6-27B INT4 OpenVINO IR (also seen with a 9B on the second card)
  • Serving args:
    --task=text_generation --target_device=GPU --plugin_config='{"ENABLE_CPU_PINNING":false}' --tool_parser qwen3coder --enable_tool_guided_generation true --reasoning_parser qwen3 --kv_cache_precision u8 --cache_size 4 --enable_prefix_caching true --cache_interval_multiplier 64 --max_num_seqs 16 --max_num_batched_tokens 4096

Reproducer (deterministic)

  1. POST a streaming chat completion ("stream": true, long generation).
  2. Read a few SSE chunks.
  3. Hard-close the client socket mid-stream (socket.close(), no graceful shutdown).
  4. Within ~90 s the server logs the assertion below and the process exits.

Observed on the very first attempt; kernel gained exactly one segfault entry.

Logs

[llm_executor][error][llm_executor.hpp:132] Error occurred in LLM executor:
Check 'm_block_table.count(seq_id) > 0' failed at
../../../../../repos/openvino.genai/src/cpp/src/continuous_batching/cache/block_manager.hpp:633

Kernel (55+ occurrences over days, all error 4 read faults):

ovms[…]: segfault at … ip …6f3 … error 4 in ovms
libstdc++.so.6.0.33[…] from mediapipe/NNNNN threads

Symbolized against the shipped binary (BuildID a1b35868566ed20139449c9674ff41b1104ec73c):
the faulting IP lands in spdlog::logger::sink_it_+0x43 — i.e. the process dies
while logging, consistent with the executor thread's error path racing the
request-drop path that has already freed the sequence's block table entry.

Impact

  • --max_num_seqs 16 means one disconnected client kills up to 15 innocent
    concurrent generations (All requests: 3; Scheduled requests: 3 logged at death).
  • On a GPU deployment the recompile/reload takes 1–6 minutes per crash.
  • Any latency spike becomes self-amplifying: slow generations → client timeouts →
    disconnects → crash → colder cache → slower generations.

Notes

  • The xe "CAT error → engine reset" events sometimes seen after the crash FOLLOW
    the segfault by 300–460 ms and are absent for many crashes — downstream cleanup,
    not the cause.
  • Not OOM (memory.events oom_kill 0), not concurrency-proportional (a second
    card at 2.3× the concurrent load crashes 51× less per stream-second — the
    discriminating variable is the mid-stream disconnect rate).
  • 2026.1 release notes mention cancellation of requests from disconnected clients
    (NPU); this looks like the GPU continuous-batching cancellation path freeing
    seq_id from m_block_table while the executor still steps the sequence.
Lingua principale
C++
Stelle
932
Fork
278
Merge medio
2g 23h
PR unite (30g)
68

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di openvinotoolkit/model_server

Tutte le issue di openvinotoolkit/model_server

Issue simili

Altre issue su C++

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.