[Bug]: /v1/responses bypasses model-owner filtering in the Ollama proxy
Maintainer thường phản hồi trong vòng 5 ngày
Đánh giá
- Độ khó
- 3/5
- Thời gian dự kiến
- Nửa ngày
- Mức phù hợp với người mới
- 84/100
Hướng nghiên cứu
Start in services/ollama-proxy/proxy.go: the inferenceEndpoints list near lines 152-165 and the handleHTTP/isInferenceRequest logic near lines 1139-1152 that populates routingModel before calling resolveCandidates. Add POST /v1/responses to the model-bearing inference routes so ownership filtering and workload accounting apply, leaving control endpoints outside filtering. Done means a regression test in the proxy's Go test suite that deliberately prefers a non-owner node, asserts it receives no /v1/responses request, and confirms the advertised owner serves it; run the package tests in services/ollama-proxy to verify.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Environment
- PAIR release v0.1.1; installed ollama-proxy reports 0.26.2 (matches the tagged versions.json).
- Two-node cluster: macOS arm64 client with an existing Ollama 0.5.12 installation; Ubuntu arm64 DGX Spark model host with PAIR-managed Ollama 0.40.0.
- Only the model host advertises
qwen3.6:35b-a3b-mtp-q4_K_M. - Codex CLI 0.160.1, Responses wire protocol.
Observed behavior
A real, isolated Codex task routed through the client's loopback PAIR proxy eventually completed and passed its eight behavior tests, but logged 32 reconnect events with:
unexpected status 404 Not Found: 404 page not found
url: http://127.0.0.1:11435/v1/responses
The same symptom occurred with a custom Responses provider and with Codex's built-in Ollama provider:
CODEX_OSS_BASE_URL=http://127.0.0.1:11435/v1 codex --oss --local-provider ollama -m qwen3.6:35b-a3b-mtp-q4_K_M
The native /api/generate route through the same proxy successfully used the remote model. The requested model is absent from the client's local engine inventory and present on the model host. I have not captured a per-request selected-node trace for the failing Responses calls; the model-filter bypass below is established directly by the tagged source.
Source-level cause
services/ollama-proxy/proxy.go at v0.1.1:
inferenceEndpointsincludes native Ollama routes and OpenAI chat/completions/embeddings, but omits/v1/responses.- In
handleHTTP,routingModelis populated only whenisInferenceRequestreturns true. resolveCandidates(routingModel)therefore receives an empty model for a Responses POST, bypassing advertised-model ownership filtering. These requests also miss inference workload accounting/reservations.
Source links:
- https://github.com/NVIDIA/Personal-AI-Router/blob/v0.1.1/services/ollama-proxy/proxy.go#L152-L165
- https://github.com/NVIDIA/Personal-AI-Router/blob/v0.1.1/services/ollama-proxy/proxy.go#L1139-L1152
This permits Responses traffic to select an engine that does not own the requested model; an older engine that lacks the Responses API returns the observed plain 404. This is distinct from the tool-call representation question in #94.
Expected behavior / regression case
Treat POST /v1/responses as model-bearing inference. With two advertised nodes and only one owning the requested model, send the request only to that owner even when scheduler priority prefers the non-owner. Record the inference workload on the serving node. Keep control endpoints outside model filtering.
A regression test should prefer a non-owner node deliberately and assert that it receives no Responses request, while the advertised owner serves the request successfully.
- Ngôn ngữ chính
- Go
- Star
- 1.6k
- Fork
- 266
- Merge trung bình
- 3 ngày 15 giờ
- Pull request đã merge (30 ngày)
- 23
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của NVIDIA/Personal-AI-Router
-
enhancement
Độ khó 4/5 Hơn một tuần Mức phù hợp với người mới 12/100
NVIDIA/Personal-AI-Router#162 ·
Maintainer thường phản hồi trong vòng 5 ngày
-
enhancement
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 25/100
NVIDIA/Personal-AI-Router#154 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 5 ngày
-
[Feature]: [Ollama] Alert the user about model pull failuresCó thể đã có người làm @ckelseynv đã nhận 1 ngày trước. Đang mởenhancement
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 56/100
NVIDIA/Personal-AI-Router#152 · 1 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 5 ngày
-
Make model download cancellation asynchronousCó thể đã có người làm @ckelseynv đã nhận 18 ngày trước. Đang mởenhancement
NVIDIA/Personal-AI-Router#115 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 5 ngày
-
Derive the model download and cancellation timeouts from measurementCó thể đã có người làm @ckelseynv đã nhận 18 ngày trước. Đang mởenhancement
NVIDIA/Personal-AI-Router#114 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 5 ngày
Tất cả issue của NVIDIA/Personal-AI-Router
Issue tương tự
-
bug frontend good first issue
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 82/100
oalders/clodhopper#133 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 62/100
peasant-labs/peasant#596 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
hatchet-dev/hatchet#5179 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 62/100
FairwindsOps/nova#484 ·