Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

RFC: Proactive Server-Side Cancellation via `Request-Timeout-Ms`

Aperta
#3,277 3 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

  • #3351 di @nightcityblade — chiusa senza merge

Valutazione

Difficoltà
3/5
Tempo stimato
1-2 giorni
Idoneità per principianti
62/100
Tipo di issue
Funzionalità
Chiarezza
Abbastanza chiara
Stato di attività
Tranquilla
Stack tecnologico
python
Ambito
api

Direzione di ricerca

Inizia in _base_client.py, in _build_headers(), e segui il modo in cui vengono rappresentati i valori di timeout del client e delle singole richieste. La modifica è completata quando un timeout wall-clock rappresentato da un semplice float produce Request-Timeout-Ms, i valori httpx.Timeout non lo producono e un header personalizzato esistente viene preservato; aggiungi o aggiorna test mirati per questi casi se viene identificata la posizione dei test circostanti.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Confirm this is a feature request for the Python library and not the underlying OpenAI API.
  • This is a feature request for the Python library
Describe the feature or improvement you're requesting

Hey, I'm writing on behalf on Baseten,

When a client times out, the server has no idea, and chaining the cancellation in services, e.g. http1 has canvats. It keeps doing expensive work (inference, token generation) that nobody will ever receive. The only signal today is a TCP disconnect — reactive, not proactive. In theory/future, this could also be adopted to e.g. issue a timeout on server side if work is unrealistic to be completed within that time.

Proposal

Send a Request-Timeout-Ms header on every request so nginx, Go contexts, and load balancers can cancel in-flight work proactively when the deadline elapses — no client disconnect needed, or cancellation chain can work proactivley. Internal services can also convert this into Request-Deadline-Ms a ms since unix epoch time, which allows server side verification in distributed systems.

Why not just use x-stainless-read-timeout?

timeout.read is a per-chunk silence threshold, not a wall-clock budget. It resets on every received chunk, so it's the wrong value to drive server-side cancellation — a healthy long running stream would get killed incorrectly.

What's a valid value?
Only a plain float timeout (e.g. OpenAI(timeout=20.0)) is a true wall-clock budget for e2e time. httpx.Timeout objects have no equivalent field. We should not send the header for those — worse than no header, as we could cancel the work on server side for this..

Proposed Implementation

# _build_headers(), _base_client.py
if "request-timeout-ms" not in lower_custom_headers:
    timeout = self.timeout if isinstance(options.timeout, NotGiven) else options.timeout
    if not isinstance(timeout, Timeout) and timeout is not None:
        headers["request-timeout-ms"] = str(int(timeout * 1000))

Prior Art

  • gRPC: grpc-timeout propagates deadlines e2e across all services — the canonical example of this pattern. Middleware can decreatse that
  • Envoy: x-envoy-upstream-rq-timeout-ms — exact same semantics, widely adopted in service meshes. Unfortunately not very cross-vendor agnostic.
  • Google Maps/Cloud API: X-Server-Timeout used for deadline propagation - unfortunately in seconds, not milliseconds.
  • Stainless SDKs: Already send x-stainless-read-timeout for observability — this builds on that foundation with correct cancellation semantics.

It would be great to have a vendor agnostic name, that could be adopted from a range of LLM projects. The stainless OpenAI API is IMO the best proxy. I think having a header we can rely on would help us save a ton of compute - i believe. Please don't make the header contain openai or stainless.

Additional context
Lingua principale
Python
Stelle
31.7k
Fork
6.8k
Merge medio
1g 4h
PR unite (30g)
126

Preparare l'ambiente

Apri in Codespaces

Avvia il container di sviluppo del progetto nel browser, con il tuo account GitHub.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di openai/openai-python

Tutte le issue di openai/openai-python

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.