Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

[Bug]: /v1/responses bypasses model-owner filtering in the Ollama proxy

Đang mở
#146 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 5 ngày

@ilaigold đang làm issue này rồi.

Từ ngày 9/10/2026.

  • #155 của @ilaigold — đang mở

Đánh giá

Độ khó
3/5
Thời gian dự kiến
Nửa ngày
Mức phù hợp với người mới
84/100
Loại issue
Lỗi
Độ rõ ràng
Đặc tả rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
go
Lĩnh vực
api, backend

Hướng nghiên cứu

Start in services/ollama-proxy/proxy.go: the inferenceEndpoints list near lines 152-165 and the handleHTTP/isInferenceRequest logic near lines 1139-1152 that populates routingModel before calling resolveCandidates. Add POST /v1/responses to the model-bearing inference routes so ownership filtering and workload accounting apply, leaving control endpoints outside filtering. Done means a regression test in the proxy's Go test suite that deliberately prefers a non-owner node, asserts it receives no /v1/responses request, and confirms the advertised owner serves it; run the package tests in services/ollama-proxy to verify.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Environment
  • PAIR release v0.1.1; installed ollama-proxy reports 0.26.2 (matches the tagged versions.json).
  • Two-node cluster: macOS arm64 client with an existing Ollama 0.5.12 installation; Ubuntu arm64 DGX Spark model host with PAIR-managed Ollama 0.40.0.
  • Only the model host advertises qwen3.6:35b-a3b-mtp-q4_K_M.
  • Codex CLI 0.160.1, Responses wire protocol.
Observed behavior

A real, isolated Codex task routed through the client's loopback PAIR proxy eventually completed and passed its eight behavior tests, but logged 32 reconnect events with:

unexpected status 404 Not Found: 404 page not found
url: http://127.0.0.1:11435/v1/responses

The same symptom occurred with a custom Responses provider and with Codex's built-in Ollama provider:

CODEX_OSS_BASE_URL=http://127.0.0.1:11435/v1 codex --oss --local-provider ollama -m qwen3.6:35b-a3b-mtp-q4_K_M

The native /api/generate route through the same proxy successfully used the remote model. The requested model is absent from the client's local engine inventory and present on the model host. I have not captured a per-request selected-node trace for the failing Responses calls; the model-filter bypass below is established directly by the tagged source.

Source-level cause

services/ollama-proxy/proxy.go at v0.1.1:

  • inferenceEndpoints includes native Ollama routes and OpenAI chat/completions/embeddings, but omits /v1/responses.
  • In handleHTTP, routingModel is populated only when isInferenceRequest returns true.
  • resolveCandidates(routingModel) therefore receives an empty model for a Responses POST, bypassing advertised-model ownership filtering. These requests also miss inference workload accounting/reservations.

Source links:

This permits Responses traffic to select an engine that does not own the requested model; an older engine that lacks the Responses API returns the observed plain 404. This is distinct from the tool-call representation question in #94.

Expected behavior / regression case

Treat POST /v1/responses as model-bearing inference. With two advertised nodes and only one owning the requested model, send the request only to that owner even when scheduler priority prefers the non-owner. Record the inference workload on the serving node. Keep control endpoints outside model filtering.

A regression test should prefer a non-owner node deliberately and assert that it receives no Responses request, while the advertised owner serves the request successfully.

Ngôn ngữ chính
Go
Star
1.6k
Fork
266
Merge trung bình
3 ngày 15 giờ
Pull request đã merge (30 ngày)
23

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của NVIDIA/Personal-AI-Router

Tất cả issue của NVIDIA/Personal-AI-Router

Issue tương tự

Thêm issue về Go

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.