RFC: Proactive Server-Side Cancellation via `Request-Timeout-Ms`
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 3/5
- Thời gian dự kiến
- 1-2 ngày
- Mức phù hợp với người mới
- 62/100
Hướng nghiên cứu
Bắt đầu trong _base_client.py tại _build_headers() và theo dõi cách biểu diễn các giá trị timeout của client và từng request. Thay đổi được hoàn tất khi một wall-clock timeout dạng float đơn giản tạo ra Request-Timeout-Ms, các giá trị httpx.Timeout không tạo ra giá trị này và một header tùy chỉnh hiện có được giữ nguyên; thêm hoặc cập nhật các test tập trung cho những trường hợp này nếu xác định được vị trí của các test liên quan.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Confirm this is a feature request for the Python library and not the underlying OpenAI API.
- This is a feature request for the Python library
Describe the feature or improvement you're requesting
Hey, I'm writing on behalf on Baseten,
When a client times out, the server has no idea, and chaining the cancellation in services, e.g. http1 has canvats. It keeps doing expensive work (inference, token generation) that nobody will ever receive. The only signal today is a TCP disconnect — reactive, not proactive. In theory/future, this could also be adopted to e.g. issue a timeout on server side if work is unrealistic to be completed within that time.
Proposal
Send a Request-Timeout-Ms header on every request so nginx, Go contexts, and load balancers can cancel in-flight work proactively when the deadline elapses — no client disconnect needed, or cancellation chain can work proactivley. Internal services can also convert this into Request-Deadline-Ms a ms since unix epoch time, which allows server side verification in distributed systems.
Why not just use x-stainless-read-timeout?
timeout.read is a per-chunk silence threshold, not a wall-clock budget. It resets on every received chunk, so it's the wrong value to drive server-side cancellation — a healthy long running stream would get killed incorrectly.
What's a valid value?
Only a plain float timeout (e.g. OpenAI(timeout=20.0)) is a true wall-clock budget for e2e time. httpx.Timeout objects have no equivalent field. We should not send the header for those — worse than no header, as we could cancel the work on server side for this..
Proposed Implementation
# _build_headers(), _base_client.py
if "request-timeout-ms" not in lower_custom_headers:
timeout = self.timeout if isinstance(options.timeout, NotGiven) else options.timeout
if not isinstance(timeout, Timeout) and timeout is not None:
headers["request-timeout-ms"] = str(int(timeout * 1000))
Prior Art
- gRPC:
grpc-timeoutpropagates deadlines e2e across all services — the canonical example of this pattern. Middleware can decreatse that - Envoy:
x-envoy-upstream-rq-timeout-ms— exact same semantics, widely adopted in service meshes. Unfortunately not very cross-vendor agnostic. - Google Maps/Cloud API:
X-Server-Timeoutused for deadline propagation - unfortunately in seconds, not milliseconds. - Stainless SDKs: Already send
x-stainless-read-timeoutfor observability — this builds on that foundation with correct cancellation semantics.
It would be great to have a vendor agnostic name, that could be adopted from a range of LLM projects. The stainless OpenAI API is IMO the best proxy. I think having a header we can rely on would help us save a ton of compute - i believe. Please don't make the header contain openai or stainless.
Additional context
- Ngôn ngữ chính
- Python
- Star
- 31.7k
- Fork
- 6.8k
- Merge trung bình
- 2 ngày 8 giờ
- Pull request đã merge (30 ngày)
- 121
Chuẩn bị môi trường
Khởi chạy dev container của dự án ngay trên trình duyệt, bằng tài khoản GitHub của bạn.
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của openai/openai-python
-
sdk-breaking-change v4
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
openai/openai-python#3837 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
openai/openai-python#3556 · 3 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
upstream
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
openai/openai-python#3294 · 3 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
openai/openai-python#2927 · 7 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug openapi
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
openai/openai-python#2502 · 7 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của openai/openai-python
Issue tương tự
-
repo-audit
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
scverse/repo-health#20 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
/context/prime scope override double-prefixes an entity-ref project and drops its scoped memoriesĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
phasespace-labs/palinode#232 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
collective/icalendar#1858 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
Maintainer thường phản hồi trong vòng 1 ngày
-
lfx-mcp cannot supply global variables: LangflowClient drops X-LANGFLOW-GLOBAL-VAR-* from envĐang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
langflow-ai/langflow#15496 ·
Maintainer thường phản hồi trong vòng 1 ngày