[Bug]: ollama-proxy: undocumented 120 s upstream response-header timeout turns a queued or cold-loading request into a 502
Maintainer antworten meist innerhalb von 5 Tagen
@ckelseynv arbeitet bereits daran.
Seit 11.9.2026.
Bewertung
Dieses Issue wurde noch nicht bewertet.
Beschreibung
Summary
ollama-proxy returns 502 {"error":"upstream error: net/http: timeout awaiting response headers"} on any upstream request that has not produced response headers within 120 s. Ollama serves one request at a time by default (OLLAMA_NUM_PARALLEL=1) and sends no headers until generation starts, so a request that sits in Ollama's queue for more than 120 s, or that waits on a model load longer than 120 s, fails through PAIR while the identical request sent directly to Ollama succeeds. The limit is not mentioned in the README, services/ollama-proxy/README.md, the flag table, or docs/troubleshooting.mdx, and there is no flag or setting that changes it. 120 s is the observed value; I did not locate the constant in the source.
PAIR version or commit
v0.1.1 release archive service-binaries-linux-x64.zip: product 0.91.7, ollama-proxy 0.26.2, nvpair-ui-broker 0.40.2, run headless via nvpair-tui 0.7.2 in tmux with NVPAIR_LOG_LEVEL=debug.
Affected component
ollama-proxy
Environment
- Node A: Ubuntu 24.04.4, kernel 6.8.0-138, RTX 3090 24 GB, Ollama 0.30.0
- Node B: Ubuntu 24.04.4, RTX 3060 12 GB, Ollama 0.17.5
- Two nodes paired. Ollama was pre-existing on both and adopted by PAIR on 11434; proxy on 11435.
Steps to reproduce
1. Queue depth (single node). Ollama on node A with qwen3.6:27b Q4_K_M (about 11 s per 200-token request on the 3090). Fire 20 independent POST /v1/chat/completions concurrently at the PAIR proxy (127.0.0.1:11435), max_tokens: 200, stream: false. Fire the same 20 at Ollama directly (127.0.0.1:11434) as the control.
| Path | Wall clock | Succeeded | Failed |
|---|---|---|---|
| Direct to Ollama | 226.5 s | 20 | 0 |
| Through PAIR | 120.3 s | 10 | 10, all 502 at 120.24 to 120.31 s |
Per-request completion times through PAIR: 12.7, 23.7, 35.2, 48.4, 58.5, 70.7, 81.5, 92.8, 104.5, 115.6 s, then ten failures at 120.2 to 120.3 s. Every request whose generation would have started after 120 s was dropped. GPU utilization on the node fell to zero within one second of the 502s: Ollama cancelled the orphaned queue entries when the proxy closed the connections.
2. Same test, 8 concurrent instead of 20, so the deepest queue position starts before 120 s:
| Path | Wall clock | Succeeded | Failed |
|---|---|---|---|
| Direct to Ollama | 86.5 s | 8 | 0 |
| Through PAIR | 85.6 s | 8 | 0 |
Per-request completions through PAIR: 11.4, 22.4, 32.6, 42.6, 52.9, 64.4, 75.1, 85.6 s. No 502s. The only variable between this run and the 10-of-20 failure is queue depth.
3. Cold load (cluster, model only on the peer). Node B holds deepseek-r1:14b (9 GB) on a slow disk; cold load takes 2 to 3 minutes. From node A, POST /v1/chat/completions for that model to node A's proxy. PAIR routes it to node B correctly, then 502s at 120 s. Node B's Ollama abandoned the load when the connection dropped (/api/ps empty afterwards). A retry 30 s later succeeded in 31.6 s because the file was by then in page cache.
Expected behavior
Either no fixed header timeout on inference routes, or a default at least as long as Ollama's own OLLAMA_LOAD_TIMEOUT (5 minutes) with a way to configure it (flag on ollama-proxy, settings.json, or a per-engine setting), plus a line in the proxy README and troubleshooting doc stating the default and the error text it produces.
Actual behavior
A silent 120 s limit that fails requests the engine would have served. Because the scheduler forwards immediately rather than queueing, the effective maximum burst size through PAIR is 120 s / per-request time per node, regardless of how many requests the client sends. For the 27B example that is about 10 requests per node. Workaround: keep bursts under that ceiling, or pre-load models with a direct keep_alive request before routing through PAIR.
Sanitized logs
One line per failed request, queue-depth case:
12:04:58.xxx [ollama-proxy] DEBUG proxy request complete id=... node_id=a9489f3b-... method=POST path=/v1/chat/completions target=127.0.0.1:11434 status=502 duration_ms=120002 ttfb_ms=0 err="net/http: timeout awaiting response headers"
Cold-load case:
11:00:20 request sent
11:02:20.115 [ollama-proxy] DEBUG proxy request complete id=65 node_id=70506f7d-... method=POST path=/v1/chat/completions target=192.168.1.32:11435 status=502 duration_ms=120002 ttfb_ms=0 err="net/http: timeout awaiting response headers"
Full request logs, the load generator, and GPU samples are available on request.
- Vorherrschende Sprache
- Go
- Sterne
- 1.6k
- Forks
- 266
- Ø Merge
- 3 T. 15 Std.
- Gemergte PRs (30 T.)
- 23
Entwicklungsumgebung
- Kein Dockerfile und keine Docker-Compose-Datei
- Hat eine Pull-Request-Vorlage
- Beitragsleitfaden lesen
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus NVIDIA/Personal-AI-Router
-
enhancement
Schwierigkeit 4/5 Über eine Woche Anfängerfreundlichkeit 12/100
NVIDIA/Personal-AI-Router#162 ·
Maintainer antworten meist innerhalb von 5 Tagen
-
enhancement
Schwierigkeit 4/5 3-5 Tage Anfängerfreundlichkeit 25/100
NVIDIA/Personal-AI-Router#154 · 1 Kommentar ·
Maintainer antworten meist innerhalb von 5 Tagen
-
[Feature]: [Ollama] Alert the user about model pull failuresEvtl. vergeben @ckelseynv hat das vor 1 Tag übernommen. Offenenhancement
Schwierigkeit 3/5 1-2 Tage Anfängerfreundlichkeit 56/100
NVIDIA/Personal-AI-Router#152 · 1 Kommentar · 1 zugewiesene Person ·
Maintainer antworten meist innerhalb von 5 Tagen
-
[Bug]: /v1/responses bypasses model-owner filtering in the Ollama proxyEvtl. vergeben @ilaigold hat das vor 2 Tagen übernommen. Offen
Schwierigkeit 3/5 Ein halber Tag Anfängerfreundlichkeit 84/100
NVIDIA/Personal-AI-Router#146 ·
Maintainer antworten meist innerhalb von 5 Tagen
-
Make model download cancellation asynchronousEvtl. vergeben @ckelseynv hat das vor 18 Tagen übernommen. Offenenhancement
NVIDIA/Personal-AI-Router#115 · 1 zugewiesene Person ·
Maintainer antworten meist innerhalb von 5 Tagen
Alle Issues in NVIDIA/Personal-AI-Router
Ähnliche Issues
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 65/100
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 73/100
Maintainer antworten meist innerhalb von 1 Tag
-
bug go
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 82/100
genkit-ai/genkit#6761 · 1 Kommentar ·
Maintainer antworten meist innerhalb von 2 Tagen
-
bug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 87/100
Maintainer antworten meist innerhalb von 2 Tagen
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
github/github-mcp-server#3475 ·
Maintainer antworten meist innerhalb von 4 Tagen