gguf_tensor_to_f16 failed when loading Qwen3.5-9B GGUF model
Los mantenedores suelen responder en 2 días
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 35/100
- Tipo de issue
- Error
- Claridad
- Bastante claro
- Estado de actividad
- Estancado
- Stack tecnológico
- cpp
- Área
- ai-infra-agents
Línea de trabajo
Comienza en src/cpp/src/gguf_utils/gguf.cpp:96, donde ocurre el fallo reportado de gguf_tensor_to_f16, y sigue la ruta de inicialización de LLM referenciada por servable_initializer.cpp. Reproduce el fallo con el comando de Windows proporcionado y el archivo GGUF de Qwen3.5; se considera completado cuando el modelo se carga y OVMS comienza a escuchar en los puertos gRPC y REST configurados.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Describe the bug
I am trying to serve the unsloth/Qwen3.5-9B-GGUF model on a baremetal Windows host using OVMS v2026.0. The server fails to start and throws a gguf_tensor_to_f16 failed error during the LLM node initialization. I suspect the GGUF parser does not yet support the tensor structure of Qwen3.5.
Since there were similar issues with other new architectures like Qwen3-VL, I would like to ask if there is a plan or timeline to support the Qwen3.5 GGUF model structure.
To Reproduce
Steps to reproduce the behavior:
- Download
Qwen3.5-9B-Q4_K_M.gguffrom Hugging Face (unsloth/Qwen3.5-9B-GGUF). - Place the file in the local directory:
C:\ovms\models\unsloth\Qwen3.5-9B-GGUF\ - Run the following OVMS launch command on a Windows baremetal host:
.\ovms.exe --source_model "unsloth/Qwen3.5-9B-GGUF" --model_repository_path \models --model_name unsloth/Qwen3.5-9B-GGUF --task text_generation --gguf_filename Qwen3.5-9B-Q4_K_M.gguf --target_device GPU --port 8000 --rest_port 9000
- See error during startup.
Expected behavior
The model should load successfully, and the OVMS server should start listening on the specified gRPC and REST ports without crashing.
Logs
[2026-03-08 14:18:36.179][22220][serving][error][servable_initializer.cpp:214] Error during llm node initialization for models_path: C:\ovms\\models\unsloth\Qwen3.5-9B-GGUF\./Qwen3.5-9B-Q4_K_M.gguf exception: Check 'data != nullptr' failed at src\cpp\src\gguf_utils\gguf.cpp:96:
[load_gguf] gguf_tensor_to_f16 failed
[2026-03-08 14:18:36.179][22220][modelmanager][error][servable_initializer.cpp:425] Error during LLM node resources initialization: The LLM Node resource initialization failed
[2026-03-08 14:18:36.179][22220][serving][error][mediapipegraphdefinition.cpp:474] Failed to process LLM node graph unsloth/Qwen3.5-9B-GGUF
[2026-03-08 14:18:36.180][22220][modelmanager][error][modelmanager.cpp:184] Couldn't start model manager
Configuration
- OVMS version:
v2026.0(OpenVINO Model Server 2026.0.0.4d3933c5, OpenVINO backend 2026.0.0) - OVMS config.json file: N/A (Using command-line parameters)
- CPU, accelerator's versions: Target device is GPU, Arc B390 with Core X7 Ultra 358H. Baremetal Windows host.
- Model repository directory structure:
C:\ovms\models\unsloth\Qwen3.5-9B-GGUF\
└── Qwen3.5-9B-Q4_K_M.gguf
- Model:
unsloth/Qwen3.5-9B-GGUFfrom Hugging Face.
Additional context
I am running this directly on Windows (baremetal), not in a Docker container. I noticed in other issues that support for newer model structures is sometimes added in later patches. Let me know if there are any workarounds for GGUF loading in the meantime.
- Lenguaje dominante
- C++
- Estrellas
- 932
- Forks
- 278
- Merge medio
- 2 d 15 h
- PR fusionados (30 d)
- 60
Preparar el entorno
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de openvinotoolkit/model_server
-
bug
Dificultad 3/5 1-2 días Aptitud para principiantes 68/100
openvinotoolkit/model_server#4609 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
`/v3/models` lists every model twice when `group_name` and `--idle_unload_timeout_seconds` are combinedPosiblemente ocupada @atobiszei la tomó hace 3 días. Abierto
openvinotoolkit/model_server#4604 · 1 asignado ·
Los mantenedores suelen responder en 2 días
-
Idle unload never happens again if the client disconnects while a sleeping graph is waking upPosiblemente ocupada @atobiszei la tomó hace 3 días. Abierto
openvinotoolkit/model_server#4603 · 1 asignado ·
Los mantenedores suelen responder en 2 días
-
bug
Dificultad 4/5 3-5 días Aptitud para principiantes 45/100
openvinotoolkit/model_server#4599 · 4 comentarios ·
Los mantenedores suelen responder en 2 días
-
Dificultad 3/5 1-2 días Aptitud para principiantes 55/100
openvinotoolkit/model_server#4586 · 1 comentario ·
Los mantenedores suelen responder en 2 días
Todos los issues de openvinotoolkit/model_server
Issues similares
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 92/100
sandialabs/seacas#945 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
ROCm/FastFlowLM#757 ·
Los mantenedores suelen responder en 1 día
-
WaterHeaterManagement: tank_percent feature reports wrong feature id (FeatureMap corruption)Abierto
Dificultad 1/5 Menos de una hora Aptitud para principiantes 92/100
espressif/esp-matter#1867 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
mltframework/shotcut#1920 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
Los mantenedores suelen responder en 1 día