[Bug]: /v1/responses bypasses model-owner filtering in the Ollama proxy
メンテナーはふだん 5 日以内に返信
評価
調査の方向性
Start in services/ollama-proxy/proxy.go: the inferenceEndpoints list near lines 152-165 and the handleHTTP/isInferenceRequest logic near lines 1139-1152 that populates routingModel before calling resolveCandidates. Add POST /v1/responses to the model-bearing inference routes so ownership filtering and workload accounting apply, leaving control endpoints outside filtering. Done means a regression test in the proxy's Go test suite that deliberately prefers a non-owner node, asserts it receives no /v1/responses request, and confirms the advertised owner serves it; run the package tests in services/ollama-proxy to verify.
索引モデルが issue の本文から書いたものです。
説明
Environment
- PAIR release v0.1.1; installed ollama-proxy reports 0.26.2 (matches the tagged versions.json).
- Two-node cluster: macOS arm64 client with an existing Ollama 0.5.12 installation; Ubuntu arm64 DGX Spark model host with PAIR-managed Ollama 0.40.0.
- Only the model host advertises
qwen3.6:35b-a3b-mtp-q4_K_M. - Codex CLI 0.160.1, Responses wire protocol.
Observed behavior
A real, isolated Codex task routed through the client's loopback PAIR proxy eventually completed and passed its eight behavior tests, but logged 32 reconnect events with:
unexpected status 404 Not Found: 404 page not found
url: http://127.0.0.1:11435/v1/responses
The same symptom occurred with a custom Responses provider and with Codex's built-in Ollama provider:
CODEX_OSS_BASE_URL=http://127.0.0.1:11435/v1 codex --oss --local-provider ollama -m qwen3.6:35b-a3b-mtp-q4_K_M
The native /api/generate route through the same proxy successfully used the remote model. The requested model is absent from the client's local engine inventory and present on the model host. I have not captured a per-request selected-node trace for the failing Responses calls; the model-filter bypass below is established directly by the tagged source.
Source-level cause
services/ollama-proxy/proxy.go at v0.1.1:
inferenceEndpointsincludes native Ollama routes and OpenAI chat/completions/embeddings, but omits/v1/responses.- In
handleHTTP,routingModelis populated only whenisInferenceRequestreturns true. resolveCandidates(routingModel)therefore receives an empty model for a Responses POST, bypassing advertised-model ownership filtering. These requests also miss inference workload accounting/reservations.
Source links:
- https://github.com/NVIDIA/Personal-AI-Router/blob/v0.1.1/services/ollama-proxy/proxy.go#L152-L165
- https://github.com/NVIDIA/Personal-AI-Router/blob/v0.1.1/services/ollama-proxy/proxy.go#L1139-L1152
This permits Responses traffic to select an engine that does not own the requested model; an older engine that lacks the Responses API returns the observed plain 404. This is distinct from the tool-call representation question in #94.
Expected behavior / regression case
Treat POST /v1/responses as model-bearing inference. With two advertised nodes and only one owning the requested model, send the request only to that owner even when scheduler priority prefers the non-owner. Record the inference workload on the serving node. Keep control endpoints outside model filtering.
A regression test should prefer a non-owner node deliberately and assert that it receives no Responses request, while the advertised owner serves the request successfully.
- 主要言語
- Go
- スター
- 1.6k
- フォーク
- 266
- 平均マージ
- 3日 15時間
- マージ済み PR(30日)
- 23
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
NVIDIA/Personal-AI-Router のほかの issue
-
enhancement
難易度 4/5 1週間以上 初心者へのやさしさ 12/100
NVIDIA/Personal-AI-Router#162 ·
メンテナーはふだん 5 日以内に返信
-
enhancement
難易度 4/5 3〜5日 初心者へのやさしさ 25/100
NVIDIA/Personal-AI-Router#154 · コメント 1 件 ·
メンテナーはふだん 5 日以内に返信
-
[Feature]: [Ollama] Alert the user about model pull failures対応中かも @ckelseynv が 1 日前に担当しました。 オープンenhancement
難易度 3/5 1〜2日 初心者へのやさしさ 56/100
NVIDIA/Personal-AI-Router#152 · コメント 1 件 · 担当者 1 名 ·
メンテナーはふだん 5 日以内に返信
-
Make model download cancellation asynchronous対応中かも @ckelseynv が 17 日前に担当しました。 オープンenhancement
NVIDIA/Personal-AI-Router#115 · 担当者 1 名 ·
メンテナーはふだん 5 日以内に返信
-
Derive the model download and cancellation timeouts from measurement対応中かも @ckelseynv が 17 日前に担当しました。 オープンenhancement
NVIDIA/Personal-AI-Router#114 · 担当者 1 名 ·
メンテナーはふだん 5 日以内に返信
NVIDIA/Personal-AI-Router の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
メンテナーはふだん 1 日以内に返信
-
bug good first issue load-balancing
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
ktrubilo9/edge-proxy#53 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
メンテナーはふだん 1 日以内に返信
-
bug needs triage receiver/dockerstats
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
open-telemetry/opentelemetry-collector-contrib#51948 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
LanternOps/breeze#8353 ·
メンテナーはふだん 1 日以内に返信