Feature request: CPU offload / HETERO:GPU,CPU support for models exceeding VRAM in continuous batching
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 35/100
- Tipo de issue
- Nueva funcionalidad
- Claridad
- Bastante claro
- Estado de actividad
- Tranquilo
- Stack tecnológico
- cpp
- Área
- machine-learning
Línea de trabajo
Comienza en pipeline_impl.cpp:169 y rastrea la validación de continuous batching que rechaza HETERO:GPU,CPU. Determina cómo se representan los dispositivos de ejecución y dónde se aplica la selección de dispositivos; el trabajo estará terminado cuando una configuración GPU-primary pueda recurrir a CPU para modelos demasiado grandes sin deshabilitar continuous batching.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
I'm running an Intel Arc Pro B50 (16 GB GDDR6) and want to serve OpenVINO/Qwen3.6-35B-A3B-int4-ov (23.1 GB) with OVMS. The model is too large to fit entirely in VRAM, so I tried --target_device HETERO:GPU,CPU to spill the overflow into system RAM (96 GB available).
OVMS rejects this at startup:
Check 'all_gpu_device || execution_devices.size() == 1' failed at pipeline_impl.cpp:169:
Continuous batching: execution device is expected to be single CPU / single GPU / multi GPUs
I understand the continuous batching pipeline currently only accepts a single device or an all-GPU HETERO config. The only workaround is --target_device CPU, which forgoes GPU acceleration entirely.
Request: Support HETERO:GPU,CPU (or an equivalent GPU-primary-with-CPU-overflow mode) so models that slightly exceed VRAM can still benefit from GPU acceleration. For context, llama.cpp's SYCL backend already handles this on the same hardware Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf runs at ~32 tok/s generation on this GPU by offloading as many layers as fit into VRAM and spilling the rest to system RAM. Having comparable functionality in OVMS would make it practical to serve mid-to-large OpenVINO models on consumer/prosumer Intel Arc GPUs without needing an exact VRAM fit.
Hardware: Intel Arc Pro B50, 16 GB GDDR6, 96 GB system RAM, LXC container on Proxmox, openvino/model_server:latest-gpu (OVMS 2026.2.0 / OpenVINO GenAI 2026.2.0.0).
- Lenguaje dominante
- C++
- Estrellas
- 932
- Forks
- 278
- Merge medio
- 3 d 5 h
- PR fusionados (30 d)
- 70
Preparar el entorno
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de openvinotoolkit/model_server
-
Dificultad 3/5 1-2 días Aptitud para principiantes 58/100
openvinotoolkit/model_server#4613 · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
enhancement
Dificultad 3/5 1-2 días Aptitud para principiantes 68/100
openvinotoolkit/model_server#4609 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
`/v3/models` lists every model twice when `group_name` and `--idle_unload_timeout_seconds` are combinedPosiblemente ocupada @atobiszei la tomó hace 4 días. Abierto
openvinotoolkit/model_server#4604 · 1 reacción · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
Idle unload never happens again if the client disconnects while a sleeping graph is waking upPosiblemente ocupada @atobiszei la tomó hace 4 días. Abierto
openvinotoolkit/model_server#4603 · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 4/5 3-5 días Aptitud para principiantes 45/100
openvinotoolkit/model_server#4599 · 4 comentarios ·
Los mantenedores suelen responder en 1 día
Todos los issues de openvinotoolkit/model_server
Issues similares
-
component: split-view platform: windows
Dificultad 2/5 1-3 horas Aptitud para principiantes 74/100
zen-browser/desktop#15616 · 1 reacción ·
Los mantenedores suelen responder en 1 día
-
area/ysql kind/bug priority/medium status/awaiting-triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
yugabyte/yugabyte-db#34415 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
WayfireWM/wayfire#3148 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
Los mantenedores suelen responder en 1 día
-
backend:DirectX
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
llvm/llvm-project#227530 ·
Los mantenedores suelen responder en 1 día