RFC: Proactive Server-Side Cancellation via `Request-Timeout-Ms`
Los mantenedores suelen responder en 1 día
@nightcityblade ya está trabajando en esto.
Desde el 1/6/2026.
- #3351 de @nightcityblade — cerrado sin fusionar
Evaluación
- Dificultad
- 3/5
- Tiempo estimado
- 1-2 días
- Aptitud para principiantes
- 62/100
Línea de trabajo
Comienza en _base_client.py, en _build_headers(), y sigue cómo se representan los valores de timeout del cliente y de cada solicitud. El cambio estará terminado cuando un timeout de reloj de pared representado por un float simple produzca Request-Timeout-Ms, los valores httpx.Timeout no lo hagan y se conserve un encabezado personalizado existente; añade o actualiza pruebas específicas para estos casos si se identifica la ubicación correspondiente entre las pruebas cercanas.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Confirm this is a feature request for the Python library and not the underlying OpenAI API.
- This is a feature request for the Python library
Describe the feature or improvement you're requesting
Hey, I'm writing on behalf on Baseten,
When a client times out, the server has no idea, and chaining the cancellation in services, e.g. http1 has canvats. It keeps doing expensive work (inference, token generation) that nobody will ever receive. The only signal today is a TCP disconnect — reactive, not proactive. In theory/future, this could also be adopted to e.g. issue a timeout on server side if work is unrealistic to be completed within that time.
Proposal
Send a Request-Timeout-Ms header on every request so nginx, Go contexts, and load balancers can cancel in-flight work proactively when the deadline elapses — no client disconnect needed, or cancellation chain can work proactivley. Internal services can also convert this into Request-Deadline-Ms a ms since unix epoch time, which allows server side verification in distributed systems.
Why not just use x-stainless-read-timeout?
timeout.read is a per-chunk silence threshold, not a wall-clock budget. It resets on every received chunk, so it's the wrong value to drive server-side cancellation — a healthy long running stream would get killed incorrectly.
What's a valid value?
Only a plain float timeout (e.g. OpenAI(timeout=20.0)) is a true wall-clock budget for e2e time. httpx.Timeout objects have no equivalent field. We should not send the header for those — worse than no header, as we could cancel the work on server side for this..
Proposed Implementation
# _build_headers(), _base_client.py
if "request-timeout-ms" not in lower_custom_headers:
timeout = self.timeout if isinstance(options.timeout, NotGiven) else options.timeout
if not isinstance(timeout, Timeout) and timeout is not None:
headers["request-timeout-ms"] = str(int(timeout * 1000))
Prior Art
- gRPC:
grpc-timeoutpropagates deadlines e2e across all services — the canonical example of this pattern. Middleware can decreatse that - Envoy:
x-envoy-upstream-rq-timeout-ms— exact same semantics, widely adopted in service meshes. Unfortunately not very cross-vendor agnostic. - Google Maps/Cloud API:
X-Server-Timeoutused for deadline propagation - unfortunately in seconds, not milliseconds. - Stainless SDKs: Already send
x-stainless-read-timeoutfor observability — this builds on that foundation with correct cancellation semantics.
It would be great to have a vendor agnostic name, that could be adopted from a range of LLM projects. The stainless OpenAI API is IMO the best proxy. I think having a header we can rely on would help us save a ton of compute - i believe. Please don't make the header contain openai or stainless.
Additional context
- Lenguaje dominante
- Python
- Estrellas
- 31.8k
- Forks
- 7.3k
- Merge medio
- 1 d 3 h
- PR fusionados (30 d)
- 131
Preparar el entorno
Inicia el contenedor de desarrollo del proyecto en tu navegador, con tu propia cuenta de GitHub.
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de openai/openai-python
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
openai/openai-python#4022 · 19 comentarios ·
Los mantenedores suelen responder en 1 día
-
fix(auth): SubjectTokenProviderError drops response and duplicates error message in workload identity providersPosiblemente ocupada @mohmedmm la tomó hace 9 días. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
openai/openai-python#4017 · 3 comentarios ·
Los mantenedores suelen responder en 1 día
-
Querystring drops explicit empty-string scalar valuesPosiblemente ocupada @sylvesterkaczmarek la tomó hace 31 días. Abiertosdk-breaking-change v4
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
openai/openai-python#3837 ·
Los mantenedores suelen responder en 1 día
-
Define + export `ServiceTiers` string literalPosiblemente ocupada @SparshGarg999 la tomó hace 57 días. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
openai/openai-python#3556 · 3 comentarios ·
Los mantenedores suelen responder en 1 día
-
Empty OPENAI_BASE_URL prevents fallback to default API endpointPosiblemente ocupada @Sehastrajit-S la tomó hace 24 días. Abiertobug
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
openai/openai-python#2927 · 6 comentarios ·
Los mantenedores suelen responder en 1 día
Todos los issues de openai/openai-python
Issues similares
-
Claiming namespace `jft63`Abiertonamespace operations
Dificultad 1/5 Menos de una hora Aptitud para principiantes 72/100
EclipseFdn/open-vsx.org#14043 ·
Los mantenedores suelen responder en 1 día
-
netbox status: needs triage type: bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
netbox-community/netbox#23376 ·
Los mantenedores suelen responder en 1 día
-
feedback simulation workshop
Dificultad 2/5 1-3 horas Aptitud para principiantes 73/100
githubnext/gh-aw-workshop#4455 ·
Los mantenedores suelen responder en 1 día
-
Triage 🩺
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
Los mantenedores suelen responder en 1 día
-
[BUG] Container scenario crashes without expected_recovery_time, kube DNS example uses retry_waitAbiertoneeds-triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 77/100
krkn-chaos/krkn#1627 · 1 comentario ·
Los mantenedores suelen responder en 1 día