Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Unload model from memory/GPU when idle for set period

オープン
#4,141 コメント 3 件 リアクション 6 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
35/100
issue の種類
機能追加
明瞭さ
おおむね明確
活発さ
静か
技術スタック
cpp
領域
ai, backend

調査の方向性

まず、関連する OpenVINO issue #33665 と、報告内でリンクされている llama.cpp のアイドル時のスリープに関するドキュメントを読み、その後、OVMS がどのようにモデルをロードして提供するかを追跡します。設定可能なアイドル動作を定義し、アイドル状態のモデルがアンロードされ、新しいリクエストが到着したときに自動的に再び利用可能になることを検証します。

索引モデルが issue の本文から書いたものです。

説明

enhancement

This issue was also reported by someone else here: https://github.com/openvinotoolkit/openvino/issues/33665

With the current implementation it is not possible to use the GPU for other workloads while OVMS is running, as it will run out of free vRAM. Other LLM servers approach this by setting a user-defined idle timeout on a model. If no API requests are made to query a specific model for a specific amount of time, that model gets removed from memory. When a new request comes in for the model, it automatically gets loaded into memory again. It would be nice if OVMS had a similar feature.

An explanation of how llama.cpp does it can be found at https://github.com/ggml-org/llama.cpp/tree/master/tools/server#sleeping-on-idle

主要言語
C++
スター
932
フォーク
278
平均マージ
2日 23時間
マージ済み PR(30日)
68

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

openvinotoolkit/model_server のほかの issue

openvinotoolkit/model_server の issue をすべて見る

似ている issue

C++ の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。