Idle unload never happens again if the client disconnects while a sleeping graph is waking up
Los mantenedores suelen responder en 1 día
@atobiszei ya está trabajando en esto.
Desde el 25/9/2026.
Evaluación
Este issue todavía no se ha evaluado.
Descripción
Describe the bug
With model groups and --idle_unload_timeout_seconds, a MediaPipe LLM graph is normally idle-unloaded after the timeout. If a client sends a request to a sleeping graph and disconnects while the graph is still waking up (reloading), the graph comes up as loaded but is never idle-unloaded again. Later normal requests don't fix this: the graph stays loaded until the server restarts.
With a 26B model this leaves ~16 GB of RAM allocated indefinitely. In our case it stayed loaded for ~4 hours with a 600 s timeout.
To Reproduce
Any small LLM works. The graph below is a standard HttpLLMCalculator graph with device: "GPU", max_num_seqs: 4, enable_prefix_caching: false, cache_size: 0.
config.json:
{"model_config_list":[{"config":{"name":"tiny","base_path":"/path/to/graph_dir","group_name":"g"}}]}
ovms --config_path config.json --rest_port 8011 --idle_unload_timeout_seconds 30 --metrics_enable
# graph starts SLEEPING; send a request and abort it during wake-up
timeout 0.5 curl -sN localhost:8011/v3/chat/completions -H 'Content-Type: application/json' \
-d '{"model":"tiny","stream":true,"messages":[{"role":"user","content":"Hi"}],"max_tokens":10}'
# wait well beyond the timeout
sleep 90
curl -s localhost:8011/metrics | grep -E '^ovms_(graph_loaded|current_graphs|requests_accepted)'
Observed
- Log:
triggering lazy wake-up reload->RELOADING->AVAILABLE->wake-up completed. After that there is noIdle unloading model groupline, ever. - Metrics stay at
ovms_graph_loaded{name="tiny"} 1,ovms_current_graphs{name="tiny"} 0, and allovms_requests_acceptedcounters are 0 (the aborted request was never counted). - Sending further normal requests afterwards and waiting beyond the timeout: still not unloaded.
Expected behavior
The graph is idle-unloaded after --idle_unload_timeout_seconds like after any other request. An aborted request during wake-up should not keep the group "in use".
Control cases (all unload correctly after the timeout)
- Normal unary request
- Streaming request aborted by the client mid-generation (graph already loaded)
My guess is that the request aborted during wake-up is never counted in ovms_requests_accepted / finished, so the group's in-use reference is never released. I haven't checked this in the code.
Configuration
- OVMS 2026.4.0.869b2186a (OpenVINO 2026.4.0-22959, GenAI 2026.4.0.0-3407), binary package for Ubuntu 24.04 running on Ubuntu 26.04
- Intel Core Ultra X7 358H (Panther Lake), Arc B390 iGPU (
device: "GPU"), 64 GB RAM - Models: gemma-4-26b-a4b-it-int4-ov (VLM continuous batching servable), reproduced with Qwen2.5-1.5B-Instruct-int4-ov
Workaround
An external watchdog reads /metrics and restarts the server if ovms_graph_loaded is 1, ovms_current_graphs is 0 and no ovms_requests_accepted counter has changed for 2× the idle timeout.
- Lenguaje dominante
- C++
- Estrellas
- 932
- Forks
- 278
- Merge medio
- 2 d 23 h
- PR fusionados (30 d)
- 68
Preparar el entorno
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de openvinotoolkit/model_server
-
Dificultad 3/5 1-2 días Aptitud para principiantes 58/100
openvinotoolkit/model_server#4613 ·
Los mantenedores suelen responder en 1 día
-
enhancement
Dificultad 3/5 1-2 días Aptitud para principiantes 68/100
openvinotoolkit/model_server#4609 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
`/v3/models` lists every model twice when `group_name` and `--idle_unload_timeout_seconds` are combinedPosiblemente ocupada @atobiszei la tomó hace 4 días. Abierto
openvinotoolkit/model_server#4604 · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 4/5 3-5 días Aptitud para principiantes 45/100
openvinotoolkit/model_server#4599 · 4 comentarios ·
Los mantenedores suelen responder en 1 día
-
Dificultad 3/5 1-2 días Aptitud para principiantes 55/100
openvinotoolkit/model_server#4586 · 1 comentario ·
Los mantenedores suelen responder en 1 día
Todos los issues de openvinotoolkit/model_server
Issues similares
-
bug
Dificultad 1/5 1-3 horas Aptitud para principiantes 88/100
isl-org/Open3D#7585 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
Unconfirmed bug
Dificultad 1/5 Menos de una hora Aptitud para principiantes 88/100
luanti-org/luanti#17605 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
area: config area: firmware priority: P2 - medium size: S type: bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
Mizithra/ActiveTerrain#16 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
grumpycoders/pcsx-redux#2171 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
Los mantenedores suelen responder en 2 días