[Bug]: Action HTTP response-header timeout is selected by a hardcoded engine name, so only Ollama gets the long client
Mantenedores costumam responder em até 5 dias
Avaliação
Esta issue ainda não foi avaliada.
Descrição
PAIR version or commit
0.91.7 (nvpair-engine-manager 0.17.4, linux/arm64). The code paths below are also present on main.
Affected component
Go services — nvpair-engine-manager
Environment
NVIDIA DGX Spark (GB10, aarch64), Ubuntu 24.04.4 LTS, driver 580.173.02, 10-node PAIR cluster. Reproduced against stock binaries with no local modifications.
Steps to reproduce
- Configure an engine other than Ollama.
- Issue an
engine:actionwhose HTTP call needs more than 30 seconds to return response headers. The realistic case is achataction against a large model with a long prompt, where prefill alone can exceed 30s before the first header is written. - Observe the action fail with a transport timeout.
Expected behavior
The per-action timeout should be a property of the action or the manifest rather than of the engine's name, so an engine author can declare that a given action may be slow.
Actual behavior
executor.go defines two clients:
engineResponseHeaderTimeout = 30 * time.Second
ollamaLoadResponseHeaderTimeout = 10 * time.Minute
and actions.go chooses between them with a literal engine-name comparison:
client := e.client
if engine == "ollama" && action == "run_model" && e.ollamaLoadClient != nil {
client = e.ollamaLoadClient
}
Every other engine therefore gets 30 seconds for every action, and no manifest field can change it.
By code inspection this is not limited to third-party engines. The shipped lmstudio.json declares a chat action (POST /v1/chat/completions) that uses the 30s client, and the vLLM manifest proposed in #9 declares a chat action as well, so a first-class vLLM engine would inherit the same ceiling.
The failure also presents badly: net/http: timeout awaiting response headers reads as a broken or hung engine rather than a client-side limit.
Sanitized logs or screenshots
action "run_model": Post "http://127.0.0.1:8800/v1/chat/completions":
net/http: timeout awaiting response headers [30s]
Scope of what I actually observed, to be precise: I hit this with an engine registered through a user manifest in <appdir>/engines/, driving engine-manager over its stdio JSON-RPC. It reproduced consistently while the model was cold and stopped once I pre-warmed the model during load, which is what isolated it to the response-header timeout rather than to the engine itself. I have not reproduced a 30s chat failure on LM Studio directly; that exposure is read from its manifest and the selection logic above.
Suggested fix
An optional per-action timeout_s in the manifest, defaulting to the current 30s, would cover this and would let the Ollama exception move into ollama.json instead of living in Go — removing the engine-name comparison entirely.
- Linguagem predominante
- Go
- Estrelas
- 1.6k
- Forks
- 266
- Merge médio
- 3d 15h
- PRs com merge (30d)
- 23
Preparar o ambiente
- Sem Dockerfile nem arquivo Docker Compose
- Tem um modelo de pull request
- Ler o guia de contribuição
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de NVIDIA/Personal-AI-Router
-
enhancement
Dificuldade 4/5 Mais de uma semana Facilidade para iniciantes 12/100
NVIDIA/Personal-AI-Router#162 ·
Mantenedores costumam responder em até 5 dias
-
enhancement
Dificuldade 4/5 3-5 dias Facilidade para iniciantes 25/100
NVIDIA/Personal-AI-Router#154 · 1 comentário ·
Mantenedores costumam responder em até 5 dias
-
[Feature]: [Ollama] Alert the user about model pull failuresTalvez já em andamento @ckelseynv assumiu há 1 dia. Abertaenhancement
Dificuldade 3/5 1-2 dias Facilidade para iniciantes 56/100
NVIDIA/Personal-AI-Router#152 · 1 comentário · 1 responsável ·
Mantenedores costumam responder em até 5 dias
-
[Bug]: /v1/responses bypasses model-owner filtering in the Ollama proxyTalvez já em andamento @ilaigold assumiu há 2 dias. Aberta
Dificuldade 3/5 Meio dia Facilidade para iniciantes 84/100
NVIDIA/Personal-AI-Router#146 ·
Mantenedores costumam responder em até 5 dias
-
Make model download cancellation asynchronousTalvez já em andamento @ckelseynv assumiu há 18 dias. Abertaenhancement
NVIDIA/Personal-AI-Router#115 · 1 responsável ·
Mantenedores costumam responder em até 5 dias
Todas as issues de NVIDIA/Personal-AI-Router
Issues semelhantes
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 76/100
prime-radiant-inc/evener#4329 ·
Mantenedores costumam responder em até 1 dia
-
Add `pdfcpu` to the pantryAberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 72/100
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 72/100
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 79/100
openwatersio/aiscast#277 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 83/100
restic/restic#22112 · 1 comentário ·
Mantenedores costumam responder em até 3 dias