[Bug]: /v1/responses bypasses model-owner filtering in the Ollama proxy
维护者通常 5 天内回复
评估
调研方向
Start in services/ollama-proxy/proxy.go: the inferenceEndpoints list near lines 152-165 and the handleHTTP/isInferenceRequest logic near lines 1139-1152 that populates routingModel before calling resolveCandidates. Add POST /v1/responses to the model-bearing inference routes so ownership filtering and workload accounting apply, leaving control endpoints outside filtering. Done means a regression test in the proxy's Go test suite that deliberately prefers a non-owner node, asserts it receives no /v1/responses request, and confirms the advertised owner serves it; run the package tests in services/ollama-proxy to verify.
由索引模型根据 Issue 内容生成。
描述
Environment
- PAIR release v0.1.1; installed ollama-proxy reports 0.26.2 (matches the tagged versions.json).
- Two-node cluster: macOS arm64 client with an existing Ollama 0.5.12 installation; Ubuntu arm64 DGX Spark model host with PAIR-managed Ollama 0.40.0.
- Only the model host advertises
qwen3.6:35b-a3b-mtp-q4_K_M. - Codex CLI 0.160.1, Responses wire protocol.
Observed behavior
A real, isolated Codex task routed through the client's loopback PAIR proxy eventually completed and passed its eight behavior tests, but logged 32 reconnect events with:
unexpected status 404 Not Found: 404 page not found
url: http://127.0.0.1:11435/v1/responses
The same symptom occurred with a custom Responses provider and with Codex's built-in Ollama provider:
CODEX_OSS_BASE_URL=http://127.0.0.1:11435/v1 codex --oss --local-provider ollama -m qwen3.6:35b-a3b-mtp-q4_K_M
The native /api/generate route through the same proxy successfully used the remote model. The requested model is absent from the client's local engine inventory and present on the model host. I have not captured a per-request selected-node trace for the failing Responses calls; the model-filter bypass below is established directly by the tagged source.
Source-level cause
services/ollama-proxy/proxy.go at v0.1.1:
inferenceEndpointsincludes native Ollama routes and OpenAI chat/completions/embeddings, but omits/v1/responses.- In
handleHTTP,routingModelis populated only whenisInferenceRequestreturns true. resolveCandidates(routingModel)therefore receives an empty model for a Responses POST, bypassing advertised-model ownership filtering. These requests also miss inference workload accounting/reservations.
Source links:
- https://github.com/NVIDIA/Personal-AI-Router/blob/v0.1.1/services/ollama-proxy/proxy.go#L152-L165
- https://github.com/NVIDIA/Personal-AI-Router/blob/v0.1.1/services/ollama-proxy/proxy.go#L1139-L1152
This permits Responses traffic to select an engine that does not own the requested model; an older engine that lacks the Responses API returns the observed plain 404. This is distinct from the tool-call representation question in #94.
Expected behavior / regression case
Treat POST /v1/responses as model-bearing inference. With two advertised nodes and only one owning the requested model, send the request only to that owner even when scheduler priority prefers the non-owner. Record the inference workload on the serving node. Keep control endpoints outside model filtering.
A regression test should prefer a non-owner node deliberately and assert that it receives no Responses request, while the advertised owner serves the request successfully.
- 主要语言
- Go
- 星标
- 1.6k
- 派生
- 266
- 平均合并
- 3 天 15 小时
- 30 天内合并 PR
- 23
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
NVIDIA/Personal-AI-Router 的其他 Issue
-
enhancement
难度 4/5 一周以上 新手友好度 12/100
NVIDIA/Personal-AI-Router#162 ·
维护者通常 5 天内回复
-
enhancement
难度 4/5 3-5 天 新手友好度 25/100
NVIDIA/Personal-AI-Router#154 · 1 条评论 ·
维护者通常 5 天内回复
-
[Feature]: [Ollama] Alert the user about model pull failures可能已有人在做 @ckelseynv 于 1 天前认领。 未关闭enhancement
难度 3/5 1-2 天 新手友好度 56/100
NVIDIA/Personal-AI-Router#152 · 1 条评论 · 已指派 1 人 ·
维护者通常 5 天内回复
-
Make model download cancellation asynchronous可能已有人在做 @ckelseynv 于 17 天前认领。 未关闭enhancement
NVIDIA/Personal-AI-Router#115 · 已指派 1 人 ·
维护者通常 5 天内回复
-
Derive the model download and cancellation timeouts from measurement可能已有人在做 @ckelseynv 于 17 天前认领。 未关闭enhancement
NVIDIA/Personal-AI-Router#114 · 已指派 1 人 ·
维护者通常 5 天内回复
查看 NVIDIA/Personal-AI-Router 的全部 Issue
相似的 Issue
-
bug frontend good first issue
难度 2/5 1-3 小时 新手友好度 86/100
维护者通常 1 天内回复
-
难度 1/5 1 小时以内 新手友好度 82/100
oalders/clodhopper#133 ·
-
难度 2/5 1-3 小时 新手友好度 62/100
peasant-labs/peasant#596 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 78/100
hatchet-dev/hatchet#5179 ·
维护者通常 1 天内回复
-
bug triage
难度 2/5 1-3 小时 新手友好度 62/100
FairwindsOps/nova#484 ·