No client-side time bound after the TCP handshake (RPC reads, Drop teardown, and TLS handshake can block forever)

Đang mở
#4 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức phù hợp với người mới
56/100
Loại issue
Lỗi
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
rust
Lĩnh vực
networking

Hướng nghiên cứu

Bắt đầu trong connection/mod.rs tại connect_stream và lần theo nơi socket đi đến tls_handshake, sau đó kiểm tra ConnectionOptions và SessionConfig. Xem xét các đường đi Drop của Session, SessionPool, SessionDataSet và PooledSession::release, cùng với các bài kiểm thử TLS hiện có. Hoàn thành khi các thao tác socket, thiết lập TLS, các lần đọc RPC và các RPC Drop best-effort đều bị giới hạn, đồng thời các trường hợp máy chủ im lặng không còn bị chặn vô thời hạn.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

bug

Summary

After TcpStream::connect_timeout succeeds, the client has no time bound on any socket operation. connect_timeout only covers the TCP handshake (connection/mod.rs connect_stream), and there is no set_read_timeout / set_write_timeout / SO_KEEPALIVE anywhere in the crate. SessionConfig::query_timeout_ms is a request-body field enforced server-side and does not cover openSession, closeSession, or the wire read itself. Every Thrift read after a successful handshake is therefore an unbounded blocking read.

This is one root cause with three manifestations (fix points in different places, one fix: a socket-level I/O timeout set before the TLS handshake and every RPC read):

F1 — ordinary RPCs hang forever

Reproduced with a local listener that completes the TCP handshake and then never replies (equivalent to a long JVM GC pause, a silently-dropping firewall, or an accepting-but-not-forwarding LB):

  • Session::open() with connect_timeout = 200ms was still blocked after 12s (the block is in the openSession read, past the connect timeout's scope).
  • With a fake server that answers openSession + requestStatementId and then goes silent, execute_non_query blocked >8s with enable_auto_reconnect = true and max_reconnect_attempts = 3. with_retry only reacts to errors, and a wedged read produces none, so auto-reconnect never fires against exactly the failure it is meant to cover.
  • Control: with the port closed, the same call returned in 176us.
F2 — Drop becomes non-recoverable

Four destructors send RPCs that wait for a response: Drop for Session (close()closeSession), Drop for SessionPool (serial entry.session.close() for every idle session), Drop for SessionDataSet (close_querycloseOperation), and the pool-closed branch of PooledSession::release. A destructor cannot return an error, be cancelled, or be given a timeout, and the failures are swallowed (let _ = / log::debug!). Reproduced against the same silent server:

  1. Session::open() succeeds, then drop(s) blocks >12s.
  2. SessionPool::new(min_size = 1) then drop(pool) blocks >12s; with 8 idle sessions only the first close is attempted and notify_all() is never reached.
  3. pool.close() followed by a guard going out of scope blocks >12s in the closed branch of release().
  4. A panic unwind through a bare Session/SessionPool never completes.

The best-effort RPC-on-drop semantics match the C#/Node SDKs and should stay; they just need the socket-level bound to become best-effort again.

F3 — TLS handshake has no deadline

tls_handshake runs rustls complete_io on a blocking socket with no read timeout. Reproduced with use_ssl = true, connect_timeout = 200ms pointed at a non-TLS IoTDB RPC port: the client sends ClientHello and waits for ServerHello; the plain framed transport misreads the record bytes as a ~369MB frame length and both sides wait forever. Session::open() was still blocked after 15s. The existing TLS tests use listeners that close the connection, so the client escapes on EOF and this shape is not covered.

Proposed fix

  • Add a socket I/O timeout (read + write) to ConnectionOptions/SessionConfig, applied to the socket immediately after connect and before the TLS handshake, so every read/write — TLS handshake, openSession, all RPCs, closeSession on drop — is bounded by it.
  • With a timeout in place, wedged reads surface as Error::Thrift, so with_retry/reconnect actually engages, and the drop paths become bounded best-effort.
Ngôn ngữ chính
Rust
Star
1
Fork
0
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của apache/iotdb-client-rust

Tất cả issue của apache/iotdb-client-rust

Issue tương tự

Thêm issue về Rust

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.