gguf_tensor_to_f16 failed when loading Qwen3.5-9B GGUF model
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 35/100
- Issue 类型
- 缺陷
- 描述清晰度
- 基本清楚
- 活跃度
- 停滞
- 技术栈
- cpp
调研方向
从 src/cpp/src/gguf_utils/gguf.cpp:96 开始,该位置会发生报告的 gguf_tensor_to_f16 失败,并跟踪 servable_initializer.cpp 引用的 LLM 初始化路径。使用提供的 Windows 命令和 Qwen3.5 GGUF 文件复现;当模型加载成功且 OVMS 开始在配置的 gRPC 和 REST 端口上监听时,即表示完成。
由索引模型根据 Issue 内容生成。
描述
Describe the bug
I am trying to serve the unsloth/Qwen3.5-9B-GGUF model on a baremetal Windows host using OVMS v2026.0. The server fails to start and throws a gguf_tensor_to_f16 failed error during the LLM node initialization. I suspect the GGUF parser does not yet support the tensor structure of Qwen3.5.
Since there were similar issues with other new architectures like Qwen3-VL, I would like to ask if there is a plan or timeline to support the Qwen3.5 GGUF model structure.
To Reproduce
Steps to reproduce the behavior:
- Download
Qwen3.5-9B-Q4_K_M.gguffrom Hugging Face (unsloth/Qwen3.5-9B-GGUF). - Place the file in the local directory:
C:\ovms\models\unsloth\Qwen3.5-9B-GGUF\ - Run the following OVMS launch command on a Windows baremetal host:
.\ovms.exe --source_model "unsloth/Qwen3.5-9B-GGUF" --model_repository_path \models --model_name unsloth/Qwen3.5-9B-GGUF --task text_generation --gguf_filename Qwen3.5-9B-Q4_K_M.gguf --target_device GPU --port 8000 --rest_port 9000
- See error during startup.
Expected behavior
The model should load successfully, and the OVMS server should start listening on the specified gRPC and REST ports without crashing.
Logs
[2026-03-08 14:18:36.179][22220][serving][error][servable_initializer.cpp:214] Error during llm node initialization for models_path: C:\ovms\\models\unsloth\Qwen3.5-9B-GGUF\./Qwen3.5-9B-Q4_K_M.gguf exception: Check 'data != nullptr' failed at src\cpp\src\gguf_utils\gguf.cpp:96:
[load_gguf] gguf_tensor_to_f16 failed
[2026-03-08 14:18:36.179][22220][modelmanager][error][servable_initializer.cpp:425] Error during LLM node resources initialization: The LLM Node resource initialization failed
[2026-03-08 14:18:36.179][22220][serving][error][mediapipegraphdefinition.cpp:474] Failed to process LLM node graph unsloth/Qwen3.5-9B-GGUF
[2026-03-08 14:18:36.180][22220][modelmanager][error][modelmanager.cpp:184] Couldn't start model manager
Configuration
- OVMS version:
v2026.0(OpenVINO Model Server 2026.0.0.4d3933c5, OpenVINO backend 2026.0.0) - OVMS config.json file: N/A (Using command-line parameters)
- CPU, accelerator's versions: Target device is GPU, Arc B390 with Core X7 Ultra 358H. Baremetal Windows host.
- Model repository directory structure:
C:\ovms\models\unsloth\Qwen3.5-9B-GGUF\
└── Qwen3.5-9B-Q4_K_M.gguf
- Model:
unsloth/Qwen3.5-9B-GGUFfrom Hugging Face.
Additional context
I am running this directly on Windows (baremetal), not in a Docker container. I noticed in other issues that support for newer model structures is sometimes added in later patches. Let me know if there are any workarounds for GGUF loading in the meantime.
- 主要语言
- C++
- 星标
- 940
- 派生
- 278
- 平均合并
- 3 天 4 小时
- 30 天内合并 PR
- 68
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
openvinotoolkit/model_server 的其他 Issue
-
难度 3/5 1-2 天 新手友好度 58/100
openvinotoolkit/model_server#4613 · 已指派 1 人 ·
维护者通常 1 天内回复
-
enhancement
难度 3/5 1-2 天 新手友好度 68/100
openvinotoolkit/model_server#4609 · 1 条评论 ·
维护者通常 1 天内回复
-
Idle unload never happens again if the client disconnects while a sleeping graph is waking up可能已有人在做 @atobiszei 于 5 天前认领。 未关闭
openvinotoolkit/model_server#4603 · 已指派 1 人 ·
维护者通常 1 天内回复
-
bug
难度 4/5 3-5 天 新手友好度 45/100
openvinotoolkit/model_server#4599 · 4 条评论 ·
维护者通常 1 天内回复
-
难度 3/5 1-2 天 新手友好度 55/100
openvinotoolkit/model_server#4586 · 5 条评论 ·
维护者通常 1 天内回复
查看 openvinotoolkit/model_server 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 75/100
mpfaffenberger/privateer_reimagined#658 ·
维护者通常 1 天内回复
-
难度 1/5 1 小时以内 新手友好度 85/100
microsoft/onnxruntime#33018 ·
维护者通常 2 天内回复
-
难度 2/5 1-3 小时 新手友好度 82/100
AXERA-TECH/ax-llm#81 ·
-
enhancement
难度 2/5 半天 新手友好度 78/100
ros-industrial/ros2_canopen#448 ·
-
难度 1/5 1 小时以内 新手友好度 78/100
维护者通常 1 天内回复