Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

[Regression] 2026.3.1-gpu hangs with 200% CPU in sched_yield / libopencl-clang on dynamic shape embeddings (working fine in 2026.3-gpu)

未关闭
#4,510 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
5/5
预计耗时
一周以上
新手友好度
35/100
Issue 类型
缺陷
描述清晰度
需要澄清
活跃度
活跃
技术栈
cpp, docker, linux

调研方向

首先,使用提供的 Docker 命令、动态 shape 的 embedding 请求、多个 --rest_workers 以及 2026.3.1-gpu 镜像复现挂起问题;将其与 2026.3-gpu 进行比较。使用提供的 GDB 回溯调查 libigdrcl.so、libopencl-clang2.so.16 和 libopenvino_intel_gpu_plugin.so。并发请求能够完成,且不会出现 deadlock、持续的 sched_yield spinning 或服务器无响应,即视为完成。

由索引模型根据 Issue 内容生成。

描述

Describe the bug
In openvino/model_server:2026.3.1-gpu, serving embedding models on Intel iGPU with multiple workers (--rest_workers > 1, e.g. 4 or 16) causes OVMS to hang indefinitely when handling variable-length texts (or concurrent dynamic shape requests).

The container CPU stays pegged at ~200% indefinitely (two threads busy-spinning at 100% each in sched_yield) and stops responding to HTTP requests.

Attaching GDB to the hanging process reveals a spinlock deadlock between libigdrcl.so (Intel OpenCL runtime) and libopencl-clang2.so.16:

  • Thread 8 (LWP 96) is stuck inside libopencl-clang2.so.16 (Compile()).
  • Worker threads (e.g. Thread 22 LWP 82) are stuck in a busy spinloop in sched_yield() inside libigdrcl.so.

To Reproduce
Steps to reproduce the behavior:

  1. Model: OpenVINO/Qwen3-Embedding-0.6B-int8-ov (or any dynamic shape embedding model).
  2. Launch OVMS on Intel GPU with multiple workers:
    docker run --device /dev/dri/renderD128:/dev/dri/renderD128 \
      -p 9200:9200 \
      openvino/model_server:2026.3.1-gpu \
      --model_repository_path /models \
      --source_model OpenVINO/Qwen3-Embedding-0.6B-int8-ov \
      --task embeddings \
      --pooling LAST \
      --target_device GPU \
      --rest_workers 4 \
      --rest_port 9200 \
      --model_name qwen3-embedding-0.6b
    
  3. Send variable-length text requests in quick succession or concurrently:
    curl -X POST http://localhost:9200/v3/embeddings \
      -H "Content-Type: application/json" \
      -d '{"model": "qwen3-embedding-0.6b", "input": "testing variable length dynamic shape sentence"}'
    
  4. OVMS hangs, container CPU pins at 200%, and requests never return.

Expected behavior
The server should properly synchronize dynamic shape OpenCL compilation without deadlocking or busy-spinning in sched_yield.

GDB Backtrace
GDB backtrace of the spinning process in 2026.3.1-gpu:

Thread 8 (LWP 96 - Compilation):

#0  0x00007fc... in ... from /usr/local/lib/libopencl-clang2.so.16
#1  0x00007fc... in Compile () from /usr/local/lib/libopencl-clang2.so.16
#2  0x00007fc... in ... from /usr/lib/x86_64-linux-gnu/intel-opencl/libigdrcl.so

Thread 22 (LWP 82 - Worker spinning in sched_yield):

#0  0x00007fc... in sched_yield () from /lib/x86_64-linux-gnu/libc.so.6
#1  0x00007fc... in ... from /usr/lib/x86_64-linux-gnu/intel-opencl/libigdrcl.so
#2  0x00007fc... in ... from /ovms/lib/libopenvino_intel_gpu_plugin.so

Environment & Hardware Configuration

  1. Host CPU: 12th Gen Intel Core i5-1235U (10 cores, 12 threads)
  2. GPU: Intel Iris Xe Graphics (Alder Lake-UP3 GT2, 80 EU, PCI ID [8086:46a8])
  3. RAM: 64 GB
  4. Host OS: Linux 6.6.x (Debian 12)
  5. Model: OpenVINO/Qwen3-Embedding-0.6B-int8-ov
  6. Docker Image: openvino/model_server:2026.3.1-gpu (Broken) vs openvino/model_server:2026.3-gpu (Working)

Root Cause Analysis & Verified Workarounds
The root trigger is that when multiple worker threads (--rest_workers > 1) handle concurrent embedding requests with variable-length text, multiple threads simultaneously call into the Intel OpenCL driver (libigdrcl.so) and trigger JIT compilation in libopencl-clang2.so.16, resulting in a spinlock deadlock (sched_yield). Furthermore, in 2026.3.1-gpu, once dynamic compilation is triggered, background compiler threads can remain permanently spinning at 200% CPU even after HTTP calls complete.

Workarounds:

  1. Reverting to openvino/model_server:2026.3-gpu with --rest_workers 1 and --truncate true: requests execute smoothly at ~40ms latency and CPU drops back to 0.07% immediately upon completion.
  2. Under 2026.3.1-gpu, setting --rest_workers 1 and --truncate true serializes requests at the HTTP layer, but libopencl-clang threads may still leak CPU loops after JIT compilation.
主要语言
C++
星标
932
派生
278
平均合并
2 天 23 小时
30 天内合并 PR
68

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

openvinotoolkit/model_server 的其他 Issue

查看 openvinotoolkit/model_server 的全部 Issue

相似的 Issue

更多 C++ Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。