Hacktoberfest 2026: die Issues, die Maintainer für den Oktober markiert haben – offen und einsteigerfreundlich. Hacktoberfest-Issues durchsuchen

[Bug]: /v1/responses bypasses model-owner filtering in the Ollama proxy

Offen
#146 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Maintainer antworten meist innerhalb von 5 Tagen

@ilaigold arbeitet bereits daran.

Seit 09.10.2026.

  • #155 von @ilaigold — offen

Bewertung

Schwierigkeit
3/5
Geschätzter Aufwand
Ein halber Tag
Anfängerfreundlichkeit
84/100
Issue-Typ
Bug
Klarheit
Klar beschrieben
Aktivitätsstatus
Aktiv
Tech-Stack
go
Bereich
api, backend

Rechercherichtung

Start in services/ollama-proxy/proxy.go: the inferenceEndpoints list near lines 152-165 and the handleHTTP/isInferenceRequest logic near lines 1139-1152 that populates routingModel before calling resolveCandidates. Add POST /v1/responses to the model-bearing inference routes so ownership filtering and workload accounting apply, leaving control endpoints outside filtering. Done means a regression test in the proxy's Go test suite that deliberately prefers a non-owner node, asserts it receives no /v1/responses request, and confirms the advertised owner serves it; run the package tests in services/ollama-proxy to verify.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Beschreibung

Environment
  • PAIR release v0.1.1; installed ollama-proxy reports 0.26.2 (matches the tagged versions.json).
  • Two-node cluster: macOS arm64 client with an existing Ollama 0.5.12 installation; Ubuntu arm64 DGX Spark model host with PAIR-managed Ollama 0.40.0.
  • Only the model host advertises qwen3.6:35b-a3b-mtp-q4_K_M.
  • Codex CLI 0.160.1, Responses wire protocol.
Observed behavior

A real, isolated Codex task routed through the client's loopback PAIR proxy eventually completed and passed its eight behavior tests, but logged 32 reconnect events with:

unexpected status 404 Not Found: 404 page not found
url: http://127.0.0.1:11435/v1/responses

The same symptom occurred with a custom Responses provider and with Codex's built-in Ollama provider:

CODEX_OSS_BASE_URL=http://127.0.0.1:11435/v1 codex --oss --local-provider ollama -m qwen3.6:35b-a3b-mtp-q4_K_M

The native /api/generate route through the same proxy successfully used the remote model. The requested model is absent from the client's local engine inventory and present on the model host. I have not captured a per-request selected-node trace for the failing Responses calls; the model-filter bypass below is established directly by the tagged source.

Source-level cause

services/ollama-proxy/proxy.go at v0.1.1:

  • inferenceEndpoints includes native Ollama routes and OpenAI chat/completions/embeddings, but omits /v1/responses.
  • In handleHTTP, routingModel is populated only when isInferenceRequest returns true.
  • resolveCandidates(routingModel) therefore receives an empty model for a Responses POST, bypassing advertised-model ownership filtering. These requests also miss inference workload accounting/reservations.

Source links:

This permits Responses traffic to select an engine that does not own the requested model; an older engine that lacks the Responses API returns the observed plain 404. This is distinct from the tool-call representation question in #94.

Expected behavior / regression case

Treat POST /v1/responses as model-bearing inference. With two advertised nodes and only one owning the requested model, send the request only to that owner even when scheduler priority prefers the non-owner. Record the inference workload on the serving node. Keep control endpoints outside model filtering.

A regression test should prefer a non-owner node deliberately and assert that it receives no Responses request, while the advertised owner serves the request successfully.

Vorherrschende Sprache
Go
Sterne
1.6k
Forks
266
Ø Merge
3 T. 15 Std.
Gemergte PRs (30 T.)
23

Entwicklungsumgebung

Erste Schritte

  1. Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
  3. Forken Sie das Repository und arbeiten Sie in einem Branch.
  4. Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.

Mehr aus NVIDIA/Personal-AI-Router

Alle Issues in NVIDIA/Personal-AI-Router

Ähnliche Issues

Weitere Issues zu Go

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.