generation_config.json path override per LLM node
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 5/5
- 预计耗时
- 一周以上
- 新手友好度
- 35/100
- Issue 类型
- 功能
- 描述清晰度
- 需要澄清
- 活跃度
- 冷清
- 技术栈
- cpp
- 领域
- ai, backend-api-design
调研方向
从 src/llm/language_model/continuous_batching/servable_initializer.cpp 开始,检查 LLMCalculatorOptions、models_path 和 mediapipe 注册如何与 graph_path 一起处理。阅读 issue 中描述的 openvino.genai ContinuousBatchingPipeline 构造和配置行为。当确定了一个双方同意的放置位置,并为每个 LLM 节点提供支持独立默认值且共享模型权重的生成配置路径时,即视为完成。
由索引模型根据 Issue 内容生成。
描述
Component: LLM continuous batching, LLMCalculatorOptions / mediapipe registration
OVMS version: 2026.1.0.72cc0624 (OpenVINO backend 2026.1.0, OpenVINO GenAI backend 2026.1.0.0)
Context
When several deployments share the same on-disk model directory but need different generation defaults (e.g. different num_assistant_tokens, temperature, or sampling settings per served endpoint), the only current option is to duplicate the model directory — including the weights — because OVMS reads generation_config.json from a fixed name inside models_path. For multi-gigabyte LLMs this is impractical.
The same problem exists for graph.pbtxt, but is already solved there: graph_path in the mediapipe config entry lets one model directory back several deployments with different graphs. There is no equivalent for generation_config.json.
Related to #4221
Question
Would it be feasible to add a per-LLM-node override for the generation-config file path — analogous to graph_path? A natural shape would be either:
- a
generation_config_pathfield inLLMCalculatorOptions(next tomodels_path), absolute or relative tomodels_path; or - a sibling field at the mediapipe config-entry level (next to
graph_path).
From a quick read of openvino.genai,ContinuousBatchingPipelineaccepts an optionalGenerationConfigat construction and exposesset_config()post-construction, so the underlying mechanism appears to be already in place. The work seems contained withinsrc/llm/language_model/continuous_batching/servable_initializer.cppon the OVMS side.
Use case
Multiple served names backed by the same model weights, each with its own generation defaults. Without per-entry generation-config selection, each variant requires a full copy of the model directory on disk.
Open questions
- Is there a reason this hasn't been exposed yet — for example, a planned different mechanism (per-deployment overrides through some other channel), or an interaction with model auto-detection/conversion that I'm missing?
- Is one of the placement options preferred from the architecture side?
- 主要语言
- C++
- 星标
- 932
- 派生
- 278
- 平均合并
- 2 天 23 小时
- 30 天内合并 PR
- 68
环境准备
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
openvinotoolkit/model_server 的其他 Issue
-
难度 3/5 1-2 天 新手友好度 58/100
openvinotoolkit/model_server#4613 ·
维护者通常 1 天内回复
-
enhancement
难度 3/5 1-2 天 新手友好度 68/100
openvinotoolkit/model_server#4609 · 1 条评论 ·
维护者通常 1 天内回复
-
`/v3/models` lists every model twice when `group_name` and `--idle_unload_timeout_seconds` are combined可能已有人在做 @atobiszei 于 4 天前认领。 未关闭
openvinotoolkit/model_server#4604 · 已指派 1 人 ·
维护者通常 1 天内回复
-
Idle unload never happens again if the client disconnects while a sleeping graph is waking up可能已有人在做 @atobiszei 于 4 天前认领。 未关闭
openvinotoolkit/model_server#4603 · 已指派 1 人 ·
维护者通常 1 天内回复
-
bug
难度 4/5 3-5 天 新手友好度 45/100
openvinotoolkit/model_server#4599 · 4 条评论 ·
维护者通常 1 天内回复
查看 openvinotoolkit/model_server 的全部 Issue
相似的 Issue
-
`enzymexla.linalg.lu` lowering fails for a tall matrix: the permutation is built with the pivot type未关闭
难度 2/5 1-3 小时 新手友好度 78/100
EnzymeAD/Enzyme-JAX#3286 ·
维护者通常 1 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 76/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 88/100
apache/iceberg-cpp#973 ·
维护者通常 1 天内回复
-
feature request
难度 1/5 1 小时以内 新手友好度 86/100
维护者通常 2 天内回复
-
难度 2/5 1-3 小时 新手友好度 84/100
维护者通常 1 天内回复