[Regression] 2026.3.1-gpu hangs with 200% CPU in sched_yield / libopencl-clang on dynamic shape embeddings (working fine in 2026.3-gpu)
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 5/5
- 预计耗时
- 一周以上
- 新手友好度
- 35/100
- Issue 类型
- 缺陷
- 描述清晰度
- 需要澄清
- 活跃度
- 活跃
- 技术栈
- cpp, docker, linux
调研方向
首先,使用提供的 Docker 命令、动态 shape 的 embedding 请求、多个 --rest_workers 以及 2026.3.1-gpu 镜像复现挂起问题;将其与 2026.3-gpu 进行比较。使用提供的 GDB 回溯调查 libigdrcl.so、libopencl-clang2.so.16 和 libopenvino_intel_gpu_plugin.so。并发请求能够完成,且不会出现 deadlock、持续的 sched_yield spinning 或服务器无响应,即视为完成。
由索引模型根据 Issue 内容生成。
描述
Describe the bug
In openvino/model_server:2026.3.1-gpu, serving embedding models on Intel iGPU with multiple workers (--rest_workers > 1, e.g. 4 or 16) causes OVMS to hang indefinitely when handling variable-length texts (or concurrent dynamic shape requests).
The container CPU stays pegged at ~200% indefinitely (two threads busy-spinning at 100% each in sched_yield) and stops responding to HTTP requests.
Attaching GDB to the hanging process reveals a spinlock deadlock between libigdrcl.so (Intel OpenCL runtime) and libopencl-clang2.so.16:
- Thread 8 (LWP 96) is stuck inside
libopencl-clang2.so.16(Compile()). - Worker threads (e.g. Thread 22 LWP 82) are stuck in a busy spinloop in
sched_yield()insidelibigdrcl.so.
To Reproduce
Steps to reproduce the behavior:
- Model:
OpenVINO/Qwen3-Embedding-0.6B-int8-ov(or any dynamic shape embedding model). - Launch OVMS on Intel GPU with multiple workers:
docker run --device /dev/dri/renderD128:/dev/dri/renderD128 \ -p 9200:9200 \ openvino/model_server:2026.3.1-gpu \ --model_repository_path /models \ --source_model OpenVINO/Qwen3-Embedding-0.6B-int8-ov \ --task embeddings \ --pooling LAST \ --target_device GPU \ --rest_workers 4 \ --rest_port 9200 \ --model_name qwen3-embedding-0.6b - Send variable-length text requests in quick succession or concurrently:
curl -X POST http://localhost:9200/v3/embeddings \ -H "Content-Type: application/json" \ -d '{"model": "qwen3-embedding-0.6b", "input": "testing variable length dynamic shape sentence"}' - OVMS hangs, container CPU pins at 200%, and requests never return.
Expected behavior
The server should properly synchronize dynamic shape OpenCL compilation without deadlocking or busy-spinning in sched_yield.
GDB Backtrace
GDB backtrace of the spinning process in 2026.3.1-gpu:
Thread 8 (LWP 96 - Compilation):
#0 0x00007fc... in ... from /usr/local/lib/libopencl-clang2.so.16
#1 0x00007fc... in Compile () from /usr/local/lib/libopencl-clang2.so.16
#2 0x00007fc... in ... from /usr/lib/x86_64-linux-gnu/intel-opencl/libigdrcl.so
Thread 22 (LWP 82 - Worker spinning in sched_yield):
#0 0x00007fc... in sched_yield () from /lib/x86_64-linux-gnu/libc.so.6
#1 0x00007fc... in ... from /usr/lib/x86_64-linux-gnu/intel-opencl/libigdrcl.so
#2 0x00007fc... in ... from /ovms/lib/libopenvino_intel_gpu_plugin.so
Environment & Hardware Configuration
- Host CPU: 12th Gen Intel Core i5-1235U (10 cores, 12 threads)
- GPU: Intel Iris Xe Graphics (Alder Lake-UP3 GT2, 80 EU, PCI ID
[8086:46a8]) - RAM: 64 GB
- Host OS: Linux 6.6.x (Debian 12)
- Model:
OpenVINO/Qwen3-Embedding-0.6B-int8-ov - Docker Image:
openvino/model_server:2026.3.1-gpu(Broken) vsopenvino/model_server:2026.3-gpu(Working)
Root Cause Analysis & Verified Workarounds
The root trigger is that when multiple worker threads (--rest_workers > 1) handle concurrent embedding requests with variable-length text, multiple threads simultaneously call into the Intel OpenCL driver (libigdrcl.so) and trigger JIT compilation in libopencl-clang2.so.16, resulting in a spinlock deadlock (sched_yield). Furthermore, in 2026.3.1-gpu, once dynamic compilation is triggered, background compiler threads can remain permanently spinning at 200% CPU even after HTTP calls complete.
Workarounds:
- Reverting to
openvino/model_server:2026.3-gpuwith--rest_workers 1and--truncate true: requests execute smoothly at ~40ms latency and CPU drops back to 0.07% immediately upon completion. - Under
2026.3.1-gpu, setting--rest_workers 1and--truncate trueserializes requests at the HTTP layer, butlibopencl-clangthreads may still leak CPU loops after JIT compilation.
- 主要语言
- C++
- 星标
- 932
- 派生
- 278
- 平均合并
- 2 天 23 小时
- 30 天内合并 PR
- 68
环境准备
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
openvinotoolkit/model_server 的其他 Issue
-
难度 3/5 1-2 天 新手友好度 58/100
openvinotoolkit/model_server#4613 ·
维护者通常 1 天内回复
-
enhancement
难度 3/5 1-2 天 新手友好度 68/100
openvinotoolkit/model_server#4609 · 1 条评论 ·
维护者通常 1 天内回复
-
`/v3/models` lists every model twice when `group_name` and `--idle_unload_timeout_seconds` are combined可能已有人在做 @atobiszei 于 4 天前认领。 未关闭
openvinotoolkit/model_server#4604 · 已指派 1 人 ·
维护者通常 1 天内回复
-
Idle unload never happens again if the client disconnects while a sleeping graph is waking up可能已有人在做 @atobiszei 于 4 天前认领。 未关闭
openvinotoolkit/model_server#4603 · 已指派 1 人 ·
维护者通常 1 天内回复
-
bug
难度 4/5 3-5 天 新手友好度 45/100
openvinotoolkit/model_server#4599 · 4 条评论 ·
维护者通常 1 天内回复
查看 openvinotoolkit/model_server 的全部 Issue
相似的 Issue
-
Unconfirmed bug
难度 1/5 1 小时以内 新手友好度 88/100
luanti-org/luanti#17605 · 1 条评论 ·
维护者通常 2 天内回复
-
area: config area: firmware priority: P2 - medium size: S type: bug
难度 2/5 1-3 小时 新手友好度 76/100
Mizithra/ActiveTerrain#16 ·
-
难度 2/5 1-3 小时 新手友好度 84/100
grumpycoders/pcsx-redux#2171 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 70/100
维护者通常 2 天内回复
-
难度 2/5 1-3 小时 新手友好度 88/100
bytedance/trae-agent#524 · 1 条评论 ·
维护者通常 1 天内回复