Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Significant Inference Time Increase with Multiple Models in OpenVINO Model Server

オープン
#3,136 コメント 6 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
28/100
issue の種類
バグ
明瞭さ
説明が足りない
活発さ
停滞
技術スタック
docker, python

調査の方向性

提供された Docker コマンドと config.json から始め、ThreadPoolExecutor クライアントを使って、単一モデルの場合と 4 モデルの場合の実行を再現します。http://localhost:811/metrics のメトリクスエンドポイントとサーバーログを比較し、設定された nireq とレイテンシの値も確認します。複数モデルでの低速化の原因を特定し、それに対処する設定変更を文書化または検証できれば完了です。

索引モデルが issue の本文から書いたものです。

説明

Significant Inference Time Increase with Multiple Models in OpenVINO Model Server

Environment

  • Operating System: Ubuntu 24.04
  • OpenVINO Version: openvino/model_server:latest (Docker container)
  • Hardware: Intel Core 12th Gen Intel(R) Core(TM) i3-1220P
  • Models: YOLOv5 models converted to OpenVINO IR format (.xml and .bin), FP32 precision
  • Deployment: Docker container with OpenVINO Model Server

Issue

I deployed the OpenVINO Model Server container with a single YOLOv5 model (FP32 precision) and observed inference times of 8-20 milliseconds per request, which is acceptable. However, when I load 4 YOLOv5 models on the same server, the inference time spikes to 30-100 milliseconds per model request. This significant increase in latency occurs despite using parallelism in my client script (via ThreadPoolExecutor) and setting "nireq": 4 per model in the server configuration.

This spike leads to higher hardware resource usage (e.g., CPU/GPU contention) and impacts real-time performance. I expected multi-model inference to maintain closer to single-model latency with proper resource allocation, especially given OpenVINO's support for parallel inference.

Logs

Single Model (model2)
[2025-03-20 17:30:41.135] Prediction duration in model model2, version 1, nireq 0: 15.680 ms
[2025-03-20 17:30:41.135] Total gRPC request processing time: 15.861 ms
[2025-03-20 17:30:41.266] Prediction duration in model model2, version 1, nireq 0: 24.077 ms
[2025-03-20 17:30:41.266] Total gRPC request processing time: 24.306 ms
[2025-03-20 17:30:41.383] Prediction duration in model model2, version 1, nireq 0: 15.227 ms
[2025-03-20 17:30:41.383] Total gRPC request processing time: 15.452 ms
Multi-Model (4 models loaded)
[2025-03-20 18:17:15.523] Prediction duration in model model1, version 1, nireq 0: 42.076 ms
[2025-03-20 18:17:15.523] Total gRPC request processing time: 42.317 ms
[2025-03-20 18:17:15.530] Prediction duration in model model2, version 1, nireq 0: 46.367 ms
[2025-03-20 18:17:15.530] Total gRPC request processing time: 46.606 ms
[2025-03-20 18:17:15.530] Prediction duration in model model3, version 1, nireq 0: 45.479 ms
[2025-03-20 18:17:15.530] Total gRPC request processing time: 45.68 ms
[2025-03-20 18:17:15.514] Prediction duration in model model4, version 1, nireq 0: 27.955 ms
[2025-03-20 18:17:15.514] Total gRPC request processing time: 28.175 ms

Configuration

  • Docker Command:
    sudo docker run -d --shm-size=23g --ulimit memlock=-1 --ulimit stack=67108864 --name openvino_model_server -v /home/ubuntu/models:/models -p 900:9000 -p 811:8000 openvino/model_server:latest --config_path /models/config.json --port 9000 --rest_port 8000 --metrics_enable --log_level DEBUG
    
    
    
  • config.json:
    {
    "model_config_list": [
      {"name": "model1", "base_path": "/models/model1", "nireq": 4, "plugin_config": {"PERFORMANCE_HINT": "LATENCY"}},
      {"name": "model2", "base_path": "/models/model2", "nireq": 4, "plugin_config": {"PERFORMANCE_HINT": "LATENCY"}},
      {"name": "model3", "base_path": "/models/model3", "nireq": 4, "plugin_config": {"PERFORMANCE_HINT": "LATENCY"}},
      {"name": "model4", "base_path": "/models/model4", "nireq": 4, "plugin_config": {"PERFORMANCE_HINT": "LATENCY"}}
    ]
    

}


Steps to Reproduce
  1. Deploy OpenVINO Model Server with a single YOLOv5 model (FP32) using the above command and a config.json containing only model2.
  2. Send gRPC inference requests (e.g., via ovmsclient) and measure latency from logs or metrics endpoint (http://localhost:811/metrics).
  3. Update config.json to include 4 YOLOv5 models (model1, model2, model3, model4).
  4. Restart the container and send parallel gRPC requests for all 4 models using a Python script with ThreadPoolExecutor.
  5. Compare inference times from logs.
Expected Behavior

With 4 models loaded and parallel inference enabled (nireq=4), I expect inference times to remain close to single-model performance (e.g., 20-30 ms total latency across all models), leveraging OpenVINO's multi-stream capabilities and parallel execution

Actual Behavior

Inference time per model increases significantly (30-100 ms per request), indicating resource contention or inefficient multi-model handling. For example, model2 jumps from 15-24 ms (single model) to 46.367 ms (multi-model).

Suggestions for optimizing resource allocation or server configuration to maintain low latency with multiple YOLOv5 models would be greatly appreciated.

Thanks in advance!

主要言語
C++
スター
932
フォーク
278
平均マージ
2日 23時間
マージ済み PR(30日)
68

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

openvinotoolkit/model_server のほかの issue

openvinotoolkit/model_server の issue をすべて見る

似ている issue

C++ の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。