Significant Inference Time Increase with Multiple Models in OpenVINO Model Server
Maintainer thường phản hồi trong vòng 2 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 28/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Cần làm rõ
- Mức độ hoạt động
- Đình trệ
- Công nghệ
- docker, python
- Lĩnh vực
- backend, devops, machine-learning
Hướng nghiên cứu
Bắt đầu với lệnh Docker và config.json được cung cấp, sau đó tái hiện các lần chạy với một mô hình và với bốn mô hình bằng client ThreadPoolExecutor. So sánh log của server với metrics endpoint tại http://localhost:811/metrics, bao gồm các giá trị nireq và độ trễ đã cấu hình. Công việc được xem là hoàn tất khi xác định được nguyên nhân gây chậm khi chạy nhiều mô hình và ghi lại hoặc xác thực một thay đổi cấu hình khắc phục vấn đề đó.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Significant Inference Time Increase with Multiple Models in OpenVINO Model Server
Environment
- Operating System: Ubuntu 24.04
- OpenVINO Version:
openvino/model_server:latest(Docker container) - Hardware: Intel Core 12th Gen Intel(R) Core(TM) i3-1220P
- Models: YOLOv5 models converted to OpenVINO IR format (
.xmland.bin), FP32 precision - Deployment: Docker container with OpenVINO Model Server
Issue
I deployed the OpenVINO Model Server container with a single YOLOv5 model (FP32 precision) and observed inference times of 8-20 milliseconds per request, which is acceptable. However, when I load 4 YOLOv5 models on the same server, the inference time spikes to 30-100 milliseconds per model request. This significant increase in latency occurs despite using parallelism in my client script (via ThreadPoolExecutor) and setting "nireq": 4 per model in the server configuration.
This spike leads to higher hardware resource usage (e.g., CPU/GPU contention) and impacts real-time performance. I expected multi-model inference to maintain closer to single-model latency with proper resource allocation, especially given OpenVINO's support for parallel inference.
Logs
Single Model (model2)
[2025-03-20 17:30:41.135] Prediction duration in model model2, version 1, nireq 0: 15.680 ms
[2025-03-20 17:30:41.135] Total gRPC request processing time: 15.861 ms
[2025-03-20 17:30:41.266] Prediction duration in model model2, version 1, nireq 0: 24.077 ms
[2025-03-20 17:30:41.266] Total gRPC request processing time: 24.306 ms
[2025-03-20 17:30:41.383] Prediction duration in model model2, version 1, nireq 0: 15.227 ms
[2025-03-20 17:30:41.383] Total gRPC request processing time: 15.452 ms
Multi-Model (4 models loaded)
[2025-03-20 18:17:15.523] Prediction duration in model model1, version 1, nireq 0: 42.076 ms
[2025-03-20 18:17:15.523] Total gRPC request processing time: 42.317 ms
[2025-03-20 18:17:15.530] Prediction duration in model model2, version 1, nireq 0: 46.367 ms
[2025-03-20 18:17:15.530] Total gRPC request processing time: 46.606 ms
[2025-03-20 18:17:15.530] Prediction duration in model model3, version 1, nireq 0: 45.479 ms
[2025-03-20 18:17:15.530] Total gRPC request processing time: 45.68 ms
[2025-03-20 18:17:15.514] Prediction duration in model model4, version 1, nireq 0: 27.955 ms
[2025-03-20 18:17:15.514] Total gRPC request processing time: 28.175 ms
Configuration
- Docker Command:
sudo docker run -d --shm-size=23g --ulimit memlock=-1 --ulimit stack=67108864 --name openvino_model_server -v /home/ubuntu/models:/models -p 900:9000 -p 811:8000 openvino/model_server:latest --config_path /models/config.json --port 9000 --rest_port 8000 --metrics_enable --log_level DEBUG - config.json:
{ "model_config_list": [ {"name": "model1", "base_path": "/models/model1", "nireq": 4, "plugin_config": {"PERFORMANCE_HINT": "LATENCY"}}, {"name": "model2", "base_path": "/models/model2", "nireq": 4, "plugin_config": {"PERFORMANCE_HINT": "LATENCY"}}, {"name": "model3", "base_path": "/models/model3", "nireq": 4, "plugin_config": {"PERFORMANCE_HINT": "LATENCY"}}, {"name": "model4", "base_path": "/models/model4", "nireq": 4, "plugin_config": {"PERFORMANCE_HINT": "LATENCY"}} ]
}
Steps to Reproduce
- Deploy OpenVINO Model Server with a single YOLOv5 model (FP32) using the above command and a
config.jsoncontaining onlymodel2. - Send gRPC inference requests (e.g., via ovmsclient) and measure latency from logs or metrics endpoint
(http://localhost:811/metrics). - Update config.json to include 4 YOLOv5 models
(model1, model2, model3, model4). - Restart the container and send parallel gRPC requests for all 4 models using a Python script with ThreadPoolExecutor.
- Compare inference times from logs.
Expected Behavior
With 4 models loaded and parallel inference enabled (nireq=4), I expect inference times to remain close to single-model performance (e.g., 20-30 ms total latency across all models), leveraging OpenVINO's multi-stream capabilities and parallel execution
Actual Behavior
Inference time per model increases significantly (30-100 ms per request), indicating resource contention or inefficient multi-model handling. For example, model2 jumps from 15-24 ms (single model) to 46.367 ms (multi-model).
Suggestions for optimizing resource allocation or server configuration to maintain low latency with multiple YOLOv5 models would be greatly appreciated.
Thanks in advance!
- Ngôn ngữ chính
- C++
- Star
- 932
- Fork
- 278
- Merge trung bình
- 2 ngày 19 giờ
- Pull request đã merge (30 ngày)
- 62
Chuẩn bị môi trường
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của openvinotoolkit/model_server
-
bug
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 68/100
openvinotoolkit/model_server#4609 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
-
`/v3/models` lists every model twice when `group_name` and `--idle_unload_timeout_seconds` are combinedCó thể đã có người làm @atobiszei đã nhận 2 ngày trước. Đang mở
openvinotoolkit/model_server#4604 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Idle unload never happens again if the client disconnects while a sleeping graph is waking upCó thể đã có người làm @atobiszei đã nhận 2 ngày trước. Đang mở
openvinotoolkit/model_server#4603 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 2 ngày
-
bug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 45/100
openvinotoolkit/model_server#4599 · 4 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 55/100
openvinotoolkit/model_server#4586 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
Tất cả issue của openvinotoolkit/model_server
Issue tương tự
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
tenstorrent/tt-metal#58057 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
maplibre/maplibre-native#4690 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
comp-query-execution
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
ClickHouse/ClickHouse#122569 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100