Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Feature request: CPU offload / HETERO:GPU,CPU support for models exceeding VRAM in continuous batching

Đang mở
#4,263 4 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
35/100
Loại issue
Tính năng
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Ít trao đổi
Công nghệ
cpp
Lĩnh vực
machine-learning

Hướng nghiên cứu

Bắt đầu tại pipeline_impl.cpp:169 và lần theo quá trình xác thực continuous batching từ chối HETERO:GPU,CPU. Xác định cách các thiết bị thực thi được biểu diễn và nơi lựa chọn thiết bị được áp đặt; hoàn thành có nghĩa là một cấu hình GPU-primary có thể chuyển sang CPU cho các mô hình quá lớn mà không vô hiệu hóa continuous batching.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

I'm running an Intel Arc Pro B50 (16 GB GDDR6) and want to serve OpenVINO/Qwen3.6-35B-A3B-int4-ov (23.1 GB) with OVMS. The model is too large to fit entirely in VRAM, so I tried --target_device HETERO:GPU,CPU to spill the overflow into system RAM (96 GB available).

OVMS rejects this at startup:

Check 'all_gpu_device || execution_devices.size() == 1' failed at pipeline_impl.cpp:169:
Continuous batching: execution device is expected to be single CPU / single GPU / multi GPUs

I understand the continuous batching pipeline currently only accepts a single device or an all-GPU HETERO config. The only workaround is --target_device CPU, which forgoes GPU acceleration entirely.

Request: Support HETERO:GPU,CPU (or an equivalent GPU-primary-with-CPU-overflow mode) so models that slightly exceed VRAM can still benefit from GPU acceleration. For context, llama.cpp's SYCL backend already handles this on the same hardware Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf runs at ~32 tok/s generation on this GPU by offloading as many layers as fit into VRAM and spilling the rest to system RAM. Having comparable functionality in OVMS would make it practical to serve mid-to-large OpenVINO models on consumer/prosumer Intel Arc GPUs without needing an exact VRAM fit.

Hardware: Intel Arc Pro B50, 16 GB GDDR6, 96 GB system RAM, LXC container on Proxmox, openvino/model_server:latest-gpu (OVMS 2026.2.0 / OpenVINO GenAI 2026.2.0.0).

Ngôn ngữ chính
C++
Star
932
Fork
278
Merge trung bình
3 ngày 5 giờ
Pull request đã merge (30 ngày)
70

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của openvinotoolkit/model_server

Tất cả issue của openvinotoolkit/model_server

Issue tương tự

Thêm issue về C++

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.