Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

generation_config.json path override per LLM node

Đang mở
#4,233 5 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 2 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
35/100
Loại issue
Tính năng
Độ rõ ràng
Cần làm rõ
Mức độ hoạt động
Ít trao đổi
Công nghệ
cpp
Lĩnh vực
ai, backend-api-design

Hướng nghiên cứu

Bắt đầu với src/llm/language_model/continuous_batching/servable_initializer.cpp và kiểm tra cách LLMCalculatorOptions, models_path và việc đăng ký mediapipe được xử lý cùng với graph_path. Đọc hành vi xây dựng và cấu hình openvino.genai ContinuousBatchingPipeline được mô tả trong issue. Công việc được xem là hoàn tất khi đã thống nhất vị trí và có một đường dẫn cấu hình sinh cho từng nút LLM, hỗ trợ các giá trị mặc định riêng biệt trong khi dùng chung trọng số mô hình.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

enhancement

Component: LLM continuous batching, LLMCalculatorOptions / mediapipe registration
OVMS version: 2026.1.0.72cc0624 (OpenVINO backend 2026.1.0, OpenVINO GenAI backend 2026.1.0.0)

Context

When several deployments share the same on-disk model directory but need different generation defaults (e.g. different num_assistant_tokens, temperature, or sampling settings per served endpoint), the only current option is to duplicate the model directory — including the weights — because OVMS reads generation_config.json from a fixed name inside models_path. For multi-gigabyte LLMs this is impractical.
The same problem exists for graph.pbtxt, but is already solved there: graph_path in the mediapipe config entry lets one model directory back several deployments with different graphs. There is no equivalent for generation_config.json.

Related to #4221

Question

Would it be feasible to add a per-LLM-node override for the generation-config file path — analogous to graph_path? A natural shape would be either:

  • a generation_config_path field in LLMCalculatorOptions (next to models_path), absolute or relative to models_path; or
  • a sibling field at the mediapipe config-entry level (next to graph_path).
    From a quick read of openvino.genai, ContinuousBatchingPipeline accepts an optional GenerationConfig at construction and exposes set_config() post-construction, so the underlying mechanism appears to be already in place. The work seems contained within src/llm/language_model/continuous_batching/servable_initializer.cpp on the OVMS side.

Use case

Multiple served names backed by the same model weights, each with its own generation defaults. Without per-entry generation-config selection, each variant requires a full copy of the model directory on disk.

Open questions

  • Is there a reason this hasn't been exposed yet — for example, a planned different mechanism (per-deployment overrides through some other channel), or an interaction with model auto-detection/conversion that I'm missing?
  • Is one of the placement options preferred from the architecture side?
Ngôn ngữ chính
C++
Star
932
Fork
278
Merge trung bình
2 ngày 15 giờ
Pull request đã merge (30 ngày)
60

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của openvinotoolkit/model_server

Tất cả issue của openvinotoolkit/model_server

Issue tương tự

Thêm issue về C++

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.