Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

[Bug]: /v1/responses bypasses model-owner filtering in the Ollama proxy

未关闭
#146 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 5 天内回复

@ilaigold 已经在做这个了。

开始于 2026年10月9日。

  • #155 来自 @ilaigold —— 未关闭

评估

难度
3/5
预计耗时
半天
新手友好度
84/100
Issue 类型
缺陷
描述清晰度
描述清楚
活跃度
活跃
技术栈
go
领域
api, backend

调研方向

Start in services/ollama-proxy/proxy.go: the inferenceEndpoints list near lines 152-165 and the handleHTTP/isInferenceRequest logic near lines 1139-1152 that populates routingModel before calling resolveCandidates. Add POST /v1/responses to the model-bearing inference routes so ownership filtering and workload accounting apply, leaving control endpoints outside filtering. Done means a regression test in the proxy's Go test suite that deliberately prefers a non-owner node, asserts it receives no /v1/responses request, and confirms the advertised owner serves it; run the package tests in services/ollama-proxy to verify.

由索引模型根据 Issue 内容生成。

描述

Environment
  • PAIR release v0.1.1; installed ollama-proxy reports 0.26.2 (matches the tagged versions.json).
  • Two-node cluster: macOS arm64 client with an existing Ollama 0.5.12 installation; Ubuntu arm64 DGX Spark model host with PAIR-managed Ollama 0.40.0.
  • Only the model host advertises qwen3.6:35b-a3b-mtp-q4_K_M.
  • Codex CLI 0.160.1, Responses wire protocol.
Observed behavior

A real, isolated Codex task routed through the client's loopback PAIR proxy eventually completed and passed its eight behavior tests, but logged 32 reconnect events with:

unexpected status 404 Not Found: 404 page not found
url: http://127.0.0.1:11435/v1/responses

The same symptom occurred with a custom Responses provider and with Codex's built-in Ollama provider:

CODEX_OSS_BASE_URL=http://127.0.0.1:11435/v1 codex --oss --local-provider ollama -m qwen3.6:35b-a3b-mtp-q4_K_M

The native /api/generate route through the same proxy successfully used the remote model. The requested model is absent from the client's local engine inventory and present on the model host. I have not captured a per-request selected-node trace for the failing Responses calls; the model-filter bypass below is established directly by the tagged source.

Source-level cause

services/ollama-proxy/proxy.go at v0.1.1:

  • inferenceEndpoints includes native Ollama routes and OpenAI chat/completions/embeddings, but omits /v1/responses.
  • In handleHTTP, routingModel is populated only when isInferenceRequest returns true.
  • resolveCandidates(routingModel) therefore receives an empty model for a Responses POST, bypassing advertised-model ownership filtering. These requests also miss inference workload accounting/reservations.

Source links:

This permits Responses traffic to select an engine that does not own the requested model; an older engine that lacks the Responses API returns the observed plain 404. This is distinct from the tool-call representation question in #94.

Expected behavior / regression case

Treat POST /v1/responses as model-bearing inference. With two advertised nodes and only one owning the requested model, send the request only to that owner even when scheduler priority prefers the non-owner. Record the inference workload on the serving node. Keep control endpoints outside model filtering.

A regression test should prefer a non-owner node deliberately and assert that it receives no Responses request, while the advertised owner serves the request successfully.

主要语言
Go
星标
1.6k
派生
266
平均合并
3 天 15 小时
30 天内合并 PR
23

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

NVIDIA/Personal-AI-Router 的其他 Issue

查看 NVIDIA/Personal-AI-Router 的全部 Issue

相似的 Issue

更多 Go Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。