Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Unload model from memory/GPU when idle for set period

未关闭
#4,141 3 条评论 6 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
5/5
预计耗时
一周以上
新手友好度
35/100
Issue 类型
功能
描述清晰度
基本清楚
活跃度
冷清
技术栈
cpp
领域
ai, backend

调研方向

首先阅读相关的 OpenVINO issue #33665 以及报告中链接的 llama.cpp 空闲休眠文档,然后追踪 OVMS 如何加载和提供模型服务。定义可配置的空闲行为,并验证处于空闲状态的模型会被卸载,并在新请求到达时自动重新可用。

由索引模型根据 Issue 内容生成。

描述

enhancement

This issue was also reported by someone else here: https://github.com/openvinotoolkit/openvino/issues/33665

With the current implementation it is not possible to use the GPU for other workloads while OVMS is running, as it will run out of free vRAM. Other LLM servers approach this by setting a user-defined idle timeout on a model. If no API requests are made to query a specific model for a specific amount of time, that model gets removed from memory. When a new request comes in for the model, it automatically gets loaded into memory again. It would be nice if OVMS had a similar feature.

An explanation of how llama.cpp does it can be found at https://github.com/ggml-org/llama.cpp/tree/master/tools/server#sleeping-on-idle

主要语言
C++
星标
932
派生
278
平均合并
3 天 5 小时
30 天内合并 PR
70

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

openvinotoolkit/model_server 的其他 Issue

查看 openvinotoolkit/model_server 的全部 Issue

相似的 Issue

更多 C++ Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。