hyper-util legacy client: an HTTP/1 request can hang forever when the connection closes while the request is being queued
Maintainer thường phản hồi trong vòng 3 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 62/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- rust
- Lĩnh vực
- backend-api-design, networking
Hướng nghiên cứu
Start at hyper_util::client::legacy::Client::try_send_request and PoolClient::poll_ready, then inspect the HTTP/1 dispatcher close path described in the issue. Use the deterministic regression tests in the linked hyper-util patch tree as the starting test reference. Done means a dispatcher close releases the held sender and the request completes with the existing retry or cancellation behavior instead of hanging.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Summary
hyper_util::client::legacy::Client can leave an HTTP/1 request future pending
forever. This happens when the pooled connection's dispatcher closes (the peer
sends RST, or a FIN on an idle keep-alive connection) at the moment the request
is being enqueued. No error is returned and the request is not retried. The
caller sees a hang that ends only when its own timeout fires.
Versions: hyper 1.9.0, hyper-util 0.1.20, tokio 1.52.3. hyper and hyper-util
master have the same code.
Mechanism
Client::try_send_requestchecks out an HTTP/1Pooled<PoolClient>and
callspooled.try_send_request(req).await. It holdspooledacross the
await.hyper::client::dispatch::Sender::try_sendcalls tokio's
UnboundedSender::send, which works in two steps:
inc_num_messages()checks the closed bit and reserves a slot, then
chan.send()publishes the value.- Suppose the requester is preempted between those two steps while the
connection task sees the peer's reset.proto::h1::dispatch::Client::recv_msg(Err)
finds no callback. It callsrx.close(), and thenrx.try_recv(). tokio's
recvreturnsPendingwhile a send is in progress (the semaphore is not
idle), andtry_recvusesnow_or_never, so it returnsNone. The
dispatcher returnsErr, and hyper-util logs
client connection error: ... ConnectionReset. On a graceful FIN, the
dispatcher finishesOkinstead and drops its receiver. - Dropping
UnboundedReceiverdrains only published values. - The requester resumes and publishes into a channel nobody will read.
Envelope::dropwould fail the callback withCanceledand give back the
request. But it runs only when the channel is dropped, and the channel
stays alive whilepooledholds its only sender.pooledis held until the
response arrives, and the response never arrives.
tokio documents that a value can be sent after the receiver is dropped
(tokio-rs/tokio#7714, tokio-rs/tokio#8278).
Evidence
In CI, a server that accepts a connection and immediately resets it
(SO_LINGER=0) produces this trace:
hyper_util::client::legacy::client: connection is ready
hyper_util::client::legacy::pool: checkout dropped
<server accepts, resets>
hyper_util::client::legacy::client: client connection error: hyper::Error(Io, Os { code: 104, kind: ConnectionReset, message: "Connection reset by peer" })
hyper_util::client::legacy::client: sending connection error to error channel
<nothing until the caller's 5 s timeout>
The client connection error line appears only when the dispatcher had no
callback and try_recv found nothing. Every other ordering fails the request
in microseconds. The hang needs a preemption inside a window of a few
instructions, so it shows up only on loaded, many-threaded runners. We saw it
five times in eight days of CI across nine variants of that test, and never in
local repetition.
Proposed fix (hyper-util)
Do not keep holding the only HTTP/1 sender after its dispatcher has closed.
While awaiting the response, try_send_request also polls
PoolClient::poll_ready for HTTP/1 connections:
Err(thewanttaker was canceled, which happens onrx.close()and on
receiver drop): captureis_reused()andconn_info, then droppooled.
That drops the last sender. The channel destructor then drops any stranded
envelope, and hyper fails the callback withCanceledplus the request. The
existingRetryablepath then retries a reused connection or returns
Canceledfor a fresh one, which is exactly what happens when the dispatcher
does see the queued request.Ok(the dispatcher asked for more work after the request was published):
it saw the request, so stop watching.
The response future is polled first on every wake, so a delivered response or
error always wins. The change does not allocate and changes no API.
The patch we currently carry against hyper-util 0.1.20 is here:
https://github.com/ferrum-edge/ferrum-edge/blob/d1136ab60f0c74221bae49cc224b265bdcb24575/docs/upstream-hyper-util-patches/001-release-h1-sender-on-dispatch-close/hyper-util-release-h1-sender-on-dispatch-close.patch
(with deterministic regression tests in the same vendored tree). We are happy to open a PR if this direction looks right.
An alternative fix in hyper: after rx.close(), keep the dispatcher polling
rx until it returns Ready(None), and cancel anything that arrives. Do the
same before dropping the receiver on a graceful shutdown. That fixes every
SendRequest user, including code that drives client::conn::http1 directly
and holds its sender while waiting. But it keeps the connection task alive
after close until the preempted sender runs again.
- Ngôn ngữ chính
- Rust
- Star
- 16.3k
- Fork
- 1.8k
- Merge trung bình
- 2 ngày 20 giờ
- Pull request đã merge (30 ngày)
- 14
Chuẩn bị môi trường
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của hyperium/hyper
-
Publicly reexport the http crateĐang mởC-feature
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 65/100
hyperium/hyper#2652 · 4 reaction ·
Maintainer thường phản hồi trong vòng 3 ngày
-
C-feature
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
hyperium/hyper#4186 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 3 ngày
-
HTTP/1 client connection is never closed or pooled when a response completes while the request body is still unsent ((Reading::KeepAlive, Writing::Body) is a terminal state)Có thể đã có người làm @BlackRabbitCoder đã nhận 12 ngày trước. Đang mởC-bug S-waiting-on-author
hyperium/hyper#4176 · 4 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 3 ngày
-
`Change graceful_shutdown function behavior` PR can cause tonic servers to hang `serve_with_incoming_shutdown` in two different casesCó thể làm lại được @seanmonstar đã nhận 30 ngày trước và không có pull request nào đang mở. Đang mởA-http2 C-bug
hyperium/hyper#4170 · 4 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 3 ngày
-
C-bug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 48/100
hyperium/hyper#4122 · 5 bình luận ·
Maintainer thường phản hồi trong vòng 3 ngày
Tất cả issue của hyperium/hyper
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
trailofbits/dylint#2107 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
area:cli bug good first issue priority:medium
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 90/100
rtk-ai/rtk#4289 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
arrays_zip with two same-named inputs fails with "ArrowArray struct has 2 children (expected 1)"Đang mởbug requires-triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
apache/datafusion-comet#6251 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
Maintainer thường phản hồi trong vòng 1 ngày
-
bug false-positive harper-core linting
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
Automattic/harper#4471 ·
Maintainer thường phản hồi trong vòng 1 ngày