Enable_prefix_caching: Suspected cache inaccuracy during prefix reuse in Linear Attention Models
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 48/100
- Issue 类型
- 缺陷
- 描述清晰度
- 基本清楚
- 活跃度
- 冷清
- 技术栈
- cpp
调研方向
检查 sequence->m_prefix_hashes 周围的前缀缓存处理,以及 KV 和 LA 缓存管理器;该 issue 指出,冲突的块大小 16 和 128,以及 get_hash() 的索引方式,可能是原因。使用启用前缀缓存的 graph.pbtxt,以及使用 40+150 和 40+80 token 用例的 pycurl.py 进行复现。当在报告的 token 范围内重复进行 temperature-0 生成时,其结果与第一次生成一致,即表示完成。
由索引模型根据 Issue 内容生成。
描述
Describe the bug
When using enable_prefix_caching in linear attention model (Qwen3.5-35B-A3B-int4-ov), if the input prompt length + output token length (total_tokens) is within range of (128~256), all subsequent generation are different even though temperature is set to 0. If total_tokens is > 256, then all generation are the same except the first generation (correct behaviour expected: all subsequent generation should be the same as the first generation). If total_tokens is <128 (no reuse since linear attention cache reuse requires >128 tokens), then all generations are the same.
To Reproduce
Steps to reproduce the behavior:
-
Steps to prepare models repository '...'
hf download <model_name> --local_dir <local_dir> -
OVMS launch command '....'
.\ovms --rest_port 8180 --config_path C:\models\Qwen3.5-35B-A3B-int4-ov\model_config.json -
Client command (additionally client code if not using official client or demo) '....'
Use simple curl command or use pycurl.py that is provided below. -
Overall step
- Enable_prefix_caching = true in graph.pbtxt.
- Curl it using input length of 40 tokens, and make it output 150 tokens (Requirement being 40+150 >= 128), using temperature 0
- Repeat this many times
- Output of all generation is different
- Control Test: Run this but with output tokens 80 (so that 40+80 <= 128), all generated output will be the same.
- Additional test: input length of 168 tokens + output tokens of 150. All generated output will be the same except the first one.
Expected behavior
All generation should be exactly the same as the first generated output regardless of the length when using temperature = 0.
Logs
OVMS logs
C:\Users\myPC\Desktop\ovms_official_2026.2.1>.\ovms --rest_port 8180 --config_path C:\models\Qwen3.5-35B-A3B-int4-ov\model_config.json
[2026-07-15 10:01:29.237][792][serving][info][server.cpp:117] OpenVINO Model Server 2026.2.1.1122f03bf
[2026-07-15 10:01:29.238][792][serving][info][server.cpp:118] OpenVINO backend 2026.2.1-21919-ede283a88e3-releases/2026/2
[2026-07-15 10:01:29.238][792][serving][info][server.cpp:121] OpenVINO GenAI backend 2026.2.1.0-3123-7dea0459b2a
[2026-07-15 10:01:29.363][792][modelmanager][info][modelmanager.cpp:180] Available devices for Open VINO: CPU, GPU, NPU
[2026-07-15 10:01:29.365][792][serving][info][capimodule.cpp:40] C-APIModule starting
[2026-07-15 10:01:29.365][792][serving][info][capimodule.cpp:42] C-APIModule started
[2026-07-15 10:01:29.365][792][serving][info][grpcservermodule.cpp:110] GRPCServerModule starting
[2026-07-15 10:01:29.365][792][serving][info][grpcservermodule.cpp:114] GRPCServerModule started
[2026-07-15 10:01:29.365][792][serving][info][grpcservermodule.cpp:115] Port was not set. GRPC server will not be started.
[2026-07-15 10:01:29.365][792][serving][info][httpservermodule.cpp:35] HTTPServerModule starting
[2026-07-15 10:01:29.365][792][serving][info][httpservermodule.cpp:39] Will start 16 REST workers
[2026-07-15 10:01:29.366][15456][serving][info][drogon_http_server.cpp:157] Binding REST server to address: 0.0.0.0:8180
[2026-07-15 10:01:29.417][792][serving][info][drogon_http_server.cpp:187] REST server listening on port 8180 with 16 unary threads and 16 streaming threads
[2026-07-15 10:01:29.418][792][serving][info][http_server.cpp:249] API key not provided via --api_key_file or API_KEY environment variable. Authentication will be disabled.
[2026-07-15 10:01:29.419][792][serving][info][httpservermodule.cpp:52] HTTPServerModule started
[2026-07-15 10:01:29.419][792][serving][info][httpservermodule.cpp:53] Started REST server at 0.0.0.0:8180
[2026-07-15 10:01:29.419][792][serving][info][servablemanagermodule.cpp:51] ServableManagerModule starting
[2026-07-15 10:01:29.421][792][serving][info][mediapipegraphdefinition.cpp:101] Graph queue globally disabled via OVMS_GRAPH_QUEUE_OFF=1 for mediapipe: Qwen3.5-35B-A3B-int4-ov
[2026-07-15 10:01:29.421][792][serving][info][mediapipegraphdefinition.cpp:479] MediapipeGraphDefinition initializing graph nodes
[2026-07-15 10:01:29.422][792][modelmanager][info][servable_initializer.cpp:450] Initializing Visual Language Model Continuous Batching servable
[2026-07-15 10:02:03.197][17036][llm_executor][info][llm_executor.hpp:128] All requests: 0; Scheduled requests: 0;
[2026-07-15 10:02:03.199][792][modelmanager][info][mediapipegraphdefinition.cpp:261] Mediapipe: Qwen3.5-35B-A3B-int4-ov inputs:
name: input; mapping: ; shape: (); precision: UNDEFINED; layout: ...
[2026-07-15 10:02:03.200][792][modelmanager][info][mediapipegraphdefinition.cpp:262] Mediapipe: Qwen3.5-35B-A3B-int4-ov outputs:
name: output; mapping: ; shape: (); precision: UNDEFINED; layout: ...
[2026-07-15 10:02:03.200][792][modelmanager][info][mediapipegraphdefinition.cpp:263] Mediapipe: Qwen3.5-35B-A3B-int4-ov kfs pass through: false
[2026-07-15 10:02:03.200][792][modelmanager][info][pipelinedefinitionstatus.hpp:60] Mediapipe: Qwen3.5-35B-A3B-int4-ov state changed to: AVAILABLE after handling: ValidationPassedEvent:
[2026-07-15 10:02:03.202][792][serving][info][servablemanagermodule.cpp:55] ServableManagerModule started
[2026-07-15 10:02:03.202][11480][modelmanager][info][modelmanager.cpp:1079] Started cleaner thread
[2026-07-15 10:02:03.202][15184][modelmanager][info][modelmanager.cpp:1060] Started model manager thread
Curling logs with comments
problem with legacy openvino
Suspect: Wrong la cache is being restored
Root cause:
1. Both kv cache manager and la cache manager is touching the data structure `sequence->m_prefix_hashes` using different block_size (kv use 16, la use 128)
2. Causes get_hash() to return a wrong hash, when input_len % block_size == 0, it return the hash by indexing the m_prefix_hashes but uses the wrong index
3. This wrong hash is used to store the la state of future la state
Test:
1. Enable_prefix_caching = true
2. Curl it using input length of 40 tokens, and make it output 150 tokens (Requirement being 40+150 >= 128), using temperature 0
3. Repeat this many times
4. Output of all generation is different
5. Control Test: Run this but with output tokens 80 (so that 40+80 <= 128), all generated output will be the same.
6. Additional test: input length of 168 tokens + output tokens of 150.
=======================================================================================
Using OVMS 2026.2.1
Scenario: When 128 < input + output token < 256, temperature = 0
Findings: All generations are different from each other.
C:\Users\myPC\Desktop\openvino.genai_2026.2.0\src\cpp\src>python pycurl.py
Success! Response from server:
{
"choices": [
{
"finish_reason": "length",
"index": 0,
"logprobs": null,
"message": {
"content": "The user is asking about OpenVINO multiple times (they repeated the question three times). This seems like they might have accidentally hit send multiple times or are emphasizing the importance of the question. I should provide a comprehensive answer about OpenVINO with examples.\n\nLet me explain what OpenVINO is:\n\n1. What it is\n2. Key features\n3. Use cases\n4. Examples with code\n5. Benefits\n\nLet me create a clear, informative response.\n</think>\n\n# What is OpenVINO?\n\n**OpenVINO\u2122** (Open Visual Inference and Neural Network Optimization) is an open-source toolkit developed by **Intel** to optimize and deploy deep learning models efficiently on Intel hardware. It enables developers to",
"role": "assistant",
"tool_calls": []
}
}
],
"created": 1784080522,
"model": "Qwen3.5-35B-A3B-int4-ov",
"object": "chat.completion",
"usage": {
"prompt_tokens": 43,
"completion_tokens": 150,
"total_tokens": 193
}
}
9:55:28.89
C:\Users\myPC\Desktop\openvino.genai_2026.2.0\src\cpp\src>python pycurl.py
Success! Response from server:
{
"choices": [
{
"finish_reason": "length",
"index": 0,
"logprobs": null,
"message": {
"content": "The user is asking about OpenVINO and wants an explanation with examples. This is a technical question about Intel's OpenVINO toolkit. Let me provide a comprehensive explanation with practical examples.\n\nOpenVINO (Open Visual Inference and Neural network Optimization) is a toolkit developed by Intel for optimizing and deploying deep learning models, particularly for computer vision applications. Let me structure this well with:\n\n1. What is OpenVINO\n2. Key features\n3. How it works\n4. Practical examples\n5. Use cases\n\nLet me create a clear, informative response with code examples.\n</think>\n\n# OpenVINO: Intel's Toolkit for AI Inference\n\n## What is OpenVINO?\n\n**OpenVINO**",
"role": "assistant",
"tool_calls": []
}
}
],
"created": 1784080540,
"model": "Qwen3.5-35B-A3B-int4-ov",
"object": "chat.completion",
"usage": {
"prompt_tokens": 43,
"completion_tokens": 150,
"total_tokens": 193
}
}
9:55:46.89
C:\Users\myPC\Desktop\openvino.genai_2026.2.0\src\cpp\src>python pycurl.py
Success! Response from server:
{
"choices": [
{
"finish_reason": "length",
"index": 0,
"logprobs": null,
"message": {
"content": "The user is asking about OpenVINO and wants an explanation with examples. Let me provide a comprehensive answer covering:\n\n1. What is OpenVINO?\n2. Key features and benefits\n3. How it works\n4. Practical examples with code\n5. Use cases\n\nLet me structure this well with clear explanations and working code examples.\n</think>\n\n# What is OpenVINO?\n\n**OpenVINO\u2122** (Open Visual Inference and Neural Network Optimization) is an open-source toolkit developed by **Intel** to optimize and accelerate deep learning inference across a wide range of hardware platforms. It enables developers to deploy AI models efficiently on CPUs, GPUs, integrated graphics, and specialized AI accelerators like Intel's Movidius",
"role": "assistant",
"tool_calls": []
}
}
],
"created": 1784080556,
"model": "Qwen3.5-35B-A3B-int4-ov",
"object": "chat.completion",
"usage": {
"prompt_tokens": 43,
"completion_tokens": 150,
"total_tokens": 193
}
}
9:56:02.44
C:\Users\myPC\Desktop\openvino.genai_2026.2.0\src\cpp\src>python pycurl.py
Success! Response from server:
{
"choices": [
{
"finish_reason": "stop",
"index": 0,
"logprobs": null,
"message": {
"content": "Thinking Process:\n\n1. **Analyze the Request:**\n * **Topic:** OpenVINO (Open Visual Inference and Neural network Optimization).\n * **Task:** Explain what it is, provide examples.\n * **Constraint:** The user's prompt seems to have a glitch/repetition at the end (\"What is OpenV OpenVINO? Please explain with examples. What is OpenVINO? Please explain with examples.",
"role": "assistant",
"tool_calls": []
}
}
],
"created": 1784080568,
"model": "Qwen3.5-35B-A3B-int4-ov",
"object": "chat.completion",
"usage": {
"prompt_tokens": 43,
"completion_tokens": 95,
"total_tokens": 138
}
}
==============================================================================
Scenario: When 0 < input + output token < 128, temperature = 0
Findings: All generations are same.
C:\Users\myPC\Desktop\openvino.genai_2026.2.0\src\cpp\src>python pycurl.py
Success! Response from server:
{
"choices": [
{
"finish_reason": "length",
"index": 0,
"logprobs": null,
"message": {
"content": "The user is asking about OpenVINO multiple times (they repeated the question three times). This seems like they might have accidentally hit send multiple times or are emphasizing the importance of the question. I should provide a comprehensive answer about OpenVINO with examples.\n\nLet me explain what OpenVINO is:\n\n1. What it is\n2. Key features\n3. Use cases\n4.",
"role": "assistant",
"tool_calls": []
}
}
],
"created": 1784080670,
"model": "Qwen3.5-35B-A3B-int4-ov",
"object": "chat.completion",
"usage": {
"prompt_tokens": 43,
"completion_tokens": 80,
"total_tokens": 123
}
}
9:57:53.97
C:\Users\myPC\Desktop\openvino.genai_2026.2.0\src\cpp\src>python pycurl.py
Success! Response from server:
{
"choices": [
{
"finish_reason": "length",
"index": 0,
"logprobs": null,
"message": {
"content": "The user is asking about OpenVINO multiple times (they repeated the question three times). This seems like they might have accidentally hit send multiple times or are emphasizing the importance of the question. I should provide a comprehensive answer about OpenVINO with examples.\n\nLet me explain what OpenVINO is:\n\n1. What it is\n2. Key features\n3. Use cases\n4.",
"role": "assistant",
"tool_calls": []
}
}
],
"created": 1784080677,
"model": "Qwen3.5-35B-A3B-int4-ov",
"object": "chat.completion",
"usage": {
"prompt_tokens": 43,
"completion_tokens": 80,
"total_tokens": 123
}
}
9:58:01.50
C:\Users\myPC\Desktop\openvino.genai_2026.2.0\src\cpp\src>python pycurl.py
Success! Response from server:
{
"choices": [
{
"finish_reason": "length",
"index": 0,
"logprobs": null,
"message": {
"content": "The user is asking about OpenVINO multiple times (they repeated the question three times). This seems like they might have accidentally hit send multiple times or are emphasizing the importance of the question. I should provide a comprehensive answer about OpenVINO with examples.\n\nLet me explain what OpenVINO is:\n\n1. What it is\n2. Key features\n3. Use cases\n4.",
"role": "assistant",
"tool_calls": []
}
}
],
"created": 1784080697,
"model": "Qwen3.5-35B-A3B-int4-ov",
"object": "chat.completion",
"usage": {
"prompt_tokens": 43,
"completion_tokens": 80,
"total_tokens": 123
}
}
==============================================================================
Scenario: When 256 < input + output token, temperature = 0
Findings: All generation are the samem except the first generation
10:01:21.39
C:\Users\myPC\Desktop\openvino.genai_2026.2.0\src\cpp\src>python pycurl.py
Success! Response from server:
{
"choices": [
{
"finish_reason": "length",
"index": 0,
"logprobs": null,
"message": {
"content": "Thinking Process:\n\n1. **Analyze the Request:**\n * The user is asking \"What is OpenVINO?\" repeatedly (20 times).\n * The core question is the same: \"What is OpenVINO? Please explain with examples.\"\n * The repetition suggests either a copy-paste error, a test of robustness, or an attempt to force a specific output format (though unlikely).\n * My task is to answer the question accurately and helpfully, ignoring the repetition (or acknowledging it subtly by providing a comprehensive answer once).\n\n2. **Define Open",
"role": "assistant",
"tool_calls": []
}
}
],
"created": 1784080929,
"model": "Qwen3.5-35B-A3B-int4-ov",
"object": "chat.completion",
"usage": {
"prompt_tokens": 230,
"completion_tokens": 128,
"total_tokens": 358
}
}
10:02:15.76
C:\Users\myPC\Desktop\openvino.genai_2026.2.0\src\cpp\src>python pycurl.py
Success! Response from server:
{
"choices": [
{
"finish_reason": "length",
"index": 0,
"logprobs": null,
"message": {
"content": "Thinking Process:\n\n1. **Analyze the Request:**\n * The user is asking \"What is OpenVINO?\" repeatedly (20 times).\n * The core question is the same: \"What is OpenVINO? Please explain with examples.\"\n * The repetition suggests either a copy-paste error, a test of robustness, or an attempt to force a specific output length/format.\n * My goal is to provide a clear, comprehensive, and accurate answer to the question \"What is OpenVINO?\" without being annoyed by the repetition. I should answer the question once",
"role": "assistant",
"tool_calls": []
}
}
],
"created": 1784080940,
"model": "Qwen3.5-35B-A3B-int4-ov",
"object": "chat.completion",
"usage": {
"prompt_tokens": 230,
"completion_tokens": 128,
"total_tokens": 358
}
}
10:02:25.98
C:\Users\myPC\Desktop\openvino.genai_2026.2.0\src\cpp\src>python pycurl.py
Success! Response from server:
{
"choices": [
{
"finish_reason": "length",
"index": 0,
"logprobs": null,
"message": {
"content": "Thinking Process:\n\n1. **Analyze the Request:**\n * The user is asking \"What is OpenVINO?\" repeatedly (20 times).\n * The core question is the same: \"What is OpenVINO? Please explain with examples.\"\n * The repetition suggests either a copy-paste error, a test of robustness, or an attempt to force a specific output length/format.\n * My goal is to provide a clear, comprehensive, and accurate answer to the question \"What is OpenVINO?\" without being annoyed by the repetition. I should answer the question once",
"role": "assistant",
"tool_calls": []
}
}
],
"created": 1784080949,
"model": "Qwen3.5-35B-A3B-int4-ov",
"object": "chat.completion",
"usage": {
"prompt_tokens": 230,
"completion_tokens": 128,
"total_tokens": 358
}
}
10:02:35.08
C:\Users\myPC\Desktop\openvino.genai_2026.2.0\src\cpp\src>python pycurl.py
Success! Response from server:
{
"choices": [
{
"finish_reason": "length",
"index": 0,
"logprobs": null,
"message": {
"content": "Thinking Process:\n\n1. **Analyze the Request:**\n * The user is asking \"What is OpenVINO?\" repeatedly (20 times).\n * The core question is the same: \"What is OpenVINO? Please explain with examples.\"\n * The repetition suggests either a copy-paste error, a test of robustness, or an attempt to force a specific output length/format.\n * My goal is to provide a clear, comprehensive, and accurate answer to the question \"What is OpenVINO?\" without being annoyed by the repetition. I should answer the question once",
"role": "assistant",
"tool_calls": []
}
}
],
"created": 1784080960,
"model": "Qwen3.5-35B-A3B-int4-ov",
"object": "chat.completion",
"usage": {
"prompt_tokens": 230,
"completion_tokens": 128,
"total_tokens": 358
}
}
Configuration
- OVMS version: 2026.2.1
- OVMS config.json file
{
"mediapipe_config_list": [
{
"name": "Qwen3.5-35B-A3B-int4-ov",
"base_path": "C:\\models\\Qwen3.5-35B-A3B-int4-ov"
}
],
"model_config_list": []
}
- CPU, accelerator's versions if applicable
Intel(R) Core(TM) Ultra X7 358H (16 CPUs), ~1.9GHz
Intel(R) Arc(TM) B390 GPU
Driver Version: 32.0.101.8826
- Model repository directory structure
C:\models\Qwen3.5-35B-A3B-int4-ov\
├── .cache\
├── .gitattributes
├── chat_template.jinja
├── config.json
├── generation_config.json
├── graph.pbtxt
├── model_config.json
├── openvino_config.json
├── openvino_detokenizer.bin
├── openvino_detokenizer.xml
├── openvino_language_model.bin
├── openvino_language_model.xml
├── openvino_text_embeddings_model.bin
├── openvino_text_embeddings_model.xml
├── openvino_text_embeddings_per_layer_model.bin
├── openvino_text_embeddings_per_layer_model.xml
├── openvino_tokenizer.bin
├── openvino_tokenizer.xml
├── openvino_vision_embeddings_model.bin
├── openvino_vision_embeddings_model.xml
├── preprocessor_config.json
├── processor_config.json
├── README.md
├── tokenizer.json
└── tokenizer_config.json
Additional context
pycurl.py
import json
import requests
# 1. Configure the API endpoint and model parameters
url = 'http://localhost:8180/v3/chat/completions'
model_name = 'Qwen3.5-35B-A3B-int4-ov'
base_sentence = 'What is OpenVINO? Please explain with examples. ' # This has 11 tokens
# Less than 1 LA block
num_repetitions = 3 # This has 10 + 11*3 = 43 tokens
prompt_content = base_sentence * num_repetitions
# 3. Construct the payload
payload = {
'model': model_name,
'messages': [
{
'role': 'user',
'content': prompt_content
}
],
'max_tokens': 150,
'temperature': 0
}
headers = {
'Content-Type': 'application/json'
}
# 4. Send the request
try:
response = requests.post(url, headers=headers, json=payload)
if response.status_code == 200:
print('Success! Response from server:')
print(json.dumps(response.json(), indent=2))
else:
print(f'Failed with status code: {response.status_code}')
print(response.text)
except requests.exceptions.RequestException as e:
print(f'An error occurred while connecting to the server: {e}')
Graph.pbtxt
input_stream: "HTTP_REQUEST_PAYLOAD:input"
output_stream: "HTTP_RESPONSE_PAYLOAD:output"
node: {
name: "LLMExecutor"
calculator: "HttpLLMCalculator"
input_stream: "LOOPBACK:loopback"
input_stream: "HTTP_REQUEST_PAYLOAD:input"
input_side_packet: "LLM_NODE_RESOURCES:llm"
output_stream: "LOOPBACK:loopback"
output_stream: "HTTP_RESPONSE_PAYLOAD:output"
input_stream_info: {
tag_index: 'LOOPBACK:0'
back_edge: true
}
node_options: {
[type.googleapis.com/mediapipe.LLMCalculatorOptions]: {
models_path: "./"
plugin_config: '{}'
enable_prefix_caching: true
dynamic_split_fuse: false
max_num_seqs: 1
max_num_batched_tokens: 131072
device: "GPU"
}
}
input_stream_handler {
input_stream_handler: "SyncSetInputStreamHandler"
options {
[mediapipe.SyncSetInputStreamHandlerOptions.ext] {
sync_set {
tag_index: "LOOPBACK:0"
}
}
}
}
}
- 主要语言
- C++
- 星标
- 940
- 派生
- 278
- 平均合并
- 3 天 4 小时
- 30 天内合并 PR
- 68
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
openvinotoolkit/model_server 的其他 Issue
-
难度 3/5 1-2 天 新手友好度 58/100
openvinotoolkit/model_server#4613 · 已指派 1 人 ·
维护者通常 1 天内回复
-
enhancement
难度 3/5 1-2 天 新手友好度 68/100
openvinotoolkit/model_server#4609 · 1 条评论 ·
维护者通常 1 天内回复
-
Idle unload never happens again if the client disconnects while a sleeping graph is waking up可能已有人在做 @atobiszei 于 6 天前认领。 未关闭
openvinotoolkit/model_server#4603 · 已指派 1 人 ·
维护者通常 1 天内回复
-
bug
难度 4/5 3-5 天 新手友好度 45/100
openvinotoolkit/model_server#4599 · 4 条评论 ·
维护者通常 1 天内回复
-
难度 3/5 1-2 天 新手友好度 55/100
openvinotoolkit/model_server#4586 · 5 条评论 ·
维护者通常 1 天内回复
查看 openvinotoolkit/model_server 的全部 Issue
相似的 Issue
-
难度 1/5 1 小时以内 新手友好度 72/100
lxqt/qtermwidget#688 ·
维护者通常 1 天内回复
-
难度 1/5 1 小时以内 新手友好度 92/100
cpinitiative/usaco-guide#6665 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 75/100
mpfaffenberger/privateer_reimagined#658 ·
维护者通常 1 天内回复
-
难度 1/5 1 小时以内 新手友好度 85/100
microsoft/onnxruntime#33018 ·
维护者通常 2 天内回复
-
Type: bug
难度 2/5 1-3 小时 新手友好度 88/100
维护者通常 2 天内回复