[Bug]: /v1/responses bypasses model-owner filtering in the Ollama proxy
Maintainer antworten meist innerhalb von 5 Tagen
Bewertung
- Schwierigkeit
- 3/5
- Geschätzter Aufwand
- Ein halber Tag
- Anfängerfreundlichkeit
- 84/100
Rechercherichtung
Start in services/ollama-proxy/proxy.go: the inferenceEndpoints list near lines 152-165 and the handleHTTP/isInferenceRequest logic near lines 1139-1152 that populates routingModel before calling resolveCandidates. Add POST /v1/responses to the model-bearing inference routes so ownership filtering and workload accounting apply, leaving control endpoints outside filtering. Done means a regression test in the proxy's Go test suite that deliberately prefers a non-owner node, asserts it receives no /v1/responses request, and confirms the advertised owner serves it; run the package tests in services/ollama-proxy to verify.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
Environment
- PAIR release v0.1.1; installed ollama-proxy reports 0.26.2 (matches the tagged versions.json).
- Two-node cluster: macOS arm64 client with an existing Ollama 0.5.12 installation; Ubuntu arm64 DGX Spark model host with PAIR-managed Ollama 0.40.0.
- Only the model host advertises
qwen3.6:35b-a3b-mtp-q4_K_M. - Codex CLI 0.160.1, Responses wire protocol.
Observed behavior
A real, isolated Codex task routed through the client's loopback PAIR proxy eventually completed and passed its eight behavior tests, but logged 32 reconnect events with:
unexpected status 404 Not Found: 404 page not found
url: http://127.0.0.1:11435/v1/responses
The same symptom occurred with a custom Responses provider and with Codex's built-in Ollama provider:
CODEX_OSS_BASE_URL=http://127.0.0.1:11435/v1 codex --oss --local-provider ollama -m qwen3.6:35b-a3b-mtp-q4_K_M
The native /api/generate route through the same proxy successfully used the remote model. The requested model is absent from the client's local engine inventory and present on the model host. I have not captured a per-request selected-node trace for the failing Responses calls; the model-filter bypass below is established directly by the tagged source.
Source-level cause
services/ollama-proxy/proxy.go at v0.1.1:
inferenceEndpointsincludes native Ollama routes and OpenAI chat/completions/embeddings, but omits/v1/responses.- In
handleHTTP,routingModelis populated only whenisInferenceRequestreturns true. resolveCandidates(routingModel)therefore receives an empty model for a Responses POST, bypassing advertised-model ownership filtering. These requests also miss inference workload accounting/reservations.
Source links:
- https://github.com/NVIDIA/Personal-AI-Router/blob/v0.1.1/services/ollama-proxy/proxy.go#L152-L165
- https://github.com/NVIDIA/Personal-AI-Router/blob/v0.1.1/services/ollama-proxy/proxy.go#L1139-L1152
This permits Responses traffic to select an engine that does not own the requested model; an older engine that lacks the Responses API returns the observed plain 404. This is distinct from the tool-call representation question in #94.
Expected behavior / regression case
Treat POST /v1/responses as model-bearing inference. With two advertised nodes and only one owning the requested model, send the request only to that owner even when scheduler priority prefers the non-owner. Record the inference workload on the serving node. Keep control endpoints outside model filtering.
A regression test should prefer a non-owner node deliberately and assert that it receives no Responses request, while the advertised owner serves the request successfully.
- Vorherrschende Sprache
- Go
- Sterne
- 1.6k
- Forks
- 266
- Ø Merge
- 3 T. 15 Std.
- Gemergte PRs (30 T.)
- 23
Entwicklungsumgebung
- Kein Dockerfile und keine Docker-Compose-Datei
- Hat eine Pull-Request-Vorlage
- Beitragsleitfaden lesen
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus NVIDIA/Personal-AI-Router
-
enhancement
Schwierigkeit 4/5 Über eine Woche Anfängerfreundlichkeit 12/100
NVIDIA/Personal-AI-Router#162 ·
Maintainer antworten meist innerhalb von 5 Tagen
-
enhancement
Schwierigkeit 4/5 3-5 Tage Anfängerfreundlichkeit 25/100
NVIDIA/Personal-AI-Router#154 · 1 Kommentar ·
Maintainer antworten meist innerhalb von 5 Tagen
-
[Feature]: [Ollama] Alert the user about model pull failuresEvtl. vergeben @ckelseynv hat das vor 2 Tagen übernommen. Offenenhancement
Schwierigkeit 3/5 1-2 Tage Anfängerfreundlichkeit 56/100
NVIDIA/Personal-AI-Router#152 · 1 Kommentar · 1 zugewiesene Person ·
Maintainer antworten meist innerhalb von 5 Tagen
-
Make model download cancellation asynchronousEvtl. vergeben @ckelseynv hat das vor 18 Tagen übernommen. Offenenhancement
NVIDIA/Personal-AI-Router#115 · 1 zugewiesene Person ·
Maintainer antworten meist innerhalb von 5 Tagen
-
Derive the model download and cancellation timeouts from measurementEvtl. vergeben @ckelseynv hat das vor 18 Tagen übernommen. Offenenhancement
NVIDIA/Personal-AI-Router#114 · 1 zugewiesene Person ·
Maintainer antworten meist innerhalb von 5 Tagen
Alle Issues in NVIDIA/Personal-AI-Router
Ähnliche Issues
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 66/100
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 76/100
prime-radiant-inc/evener#4329 ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 72/100
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 72/100
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 79/100
openwatersio/aiscast#277 ·
Maintainer antworten meist innerhalb von 1 Tag