No client-side time bound after the TCP handshake (RPC reads, Drop teardown, and TLS handshake can block forever)
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 56/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- rust
- Domain
- networking
Research direction
Start in connection/mod.rs at connect_stream and trace where the socket reaches tls_handshake, then inspect ConnectionOptions and SessionConfig. Review the Drop paths for Session, SessionPool, SessionDataSet, and PooledSession::release, along with the existing TLS tests. Done means socket operations, TLS setup, RPC reads, and best-effort drop RPCs are bounded and the silent-server cases no longer block indefinitely.
Written by the indexing model from the issue text.
Description
Summary
After TcpStream::connect_timeout succeeds, the client has no time bound on any socket operation. connect_timeout only covers the TCP handshake (connection/mod.rs connect_stream), and there is no set_read_timeout / set_write_timeout / SO_KEEPALIVE anywhere in the crate. SessionConfig::query_timeout_ms is a request-body field enforced server-side and does not cover openSession, closeSession, or the wire read itself. Every Thrift read after a successful handshake is therefore an unbounded blocking read.
This is one root cause with three manifestations (fix points in different places, one fix: a socket-level I/O timeout set before the TLS handshake and every RPC read):
F1 — ordinary RPCs hang forever
Reproduced with a local listener that completes the TCP handshake and then never replies (equivalent to a long JVM GC pause, a silently-dropping firewall, or an accepting-but-not-forwarding LB):
Session::open()withconnect_timeout = 200mswas still blocked after 12s (the block is in theopenSessionread, past the connect timeout's scope).- With a fake server that answers
openSession+requestStatementIdand then goes silent,execute_non_queryblocked >8s withenable_auto_reconnect = trueandmax_reconnect_attempts = 3.with_retryonly reacts to errors, and a wedged read produces none, so auto-reconnect never fires against exactly the failure it is meant to cover. - Control: with the port closed, the same call returned in 176us.
F2 — Drop becomes non-recoverable
Four destructors send RPCs that wait for a response: Drop for Session (close() → closeSession), Drop for SessionPool (serial entry.session.close() for every idle session), Drop for SessionDataSet (close_query → closeOperation), and the pool-closed branch of PooledSession::release. A destructor cannot return an error, be cancelled, or be given a timeout, and the failures are swallowed (let _ = / log::debug!). Reproduced against the same silent server:
Session::open()succeeds, thendrop(s)blocks >12s.SessionPool::new(min_size = 1)thendrop(pool)blocks >12s; with 8 idle sessions only the first close is attempted andnotify_all()is never reached.pool.close()followed by a guard going out of scope blocks >12s in the closed branch ofrelease().- A panic unwind through a bare
Session/SessionPoolnever completes.
The best-effort RPC-on-drop semantics match the C#/Node SDKs and should stay; they just need the socket-level bound to become best-effort again.
F3 — TLS handshake has no deadline
tls_handshake runs rustls complete_io on a blocking socket with no read timeout. Reproduced with use_ssl = true, connect_timeout = 200ms pointed at a non-TLS IoTDB RPC port: the client sends ClientHello and waits for ServerHello; the plain framed transport misreads the record bytes as a ~369MB frame length and both sides wait forever. Session::open() was still blocked after 15s. The existing TLS tests use listeners that close the connection, so the client escapes on EOF and this shape is not covered.
Proposed fix
- Add a socket I/O timeout (read + write) to
ConnectionOptions/SessionConfig, applied to the socket immediately after connect and before the TLS handshake, so every read/write — TLS handshake,openSession, all RPCs,closeSessionon drop — is bounded by it. - With a timeout in place, wedged reads surface as
Error::Thrift, sowith_retry/reconnect actually engages, and the drop paths become bounded best-effort.
- Dominant language
- Rust
- Stars
- 1
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from apache/iotdb-client-rust
-
Session::open does not fail over when openSession/requestStatementId fails (only TCP-layer failover) Openbug
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
-
enhancement
Difficulty 3/5 1-2 days Newbie friendliness 65/100
apache/iotdb-client-rust#10 ·
-
enhancement
Difficulty 4/5 3-5 days Newbie friendliness 45/100
-
bug
Difficulty 4/5 3-5 days Newbie friendliness 55/100
-
bug
Difficulty 3/5 1-2 days Newbie friendliness 68/100
All issues in apache/iotdb-client-rust
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Eynzof/Hermes-CN-Desktop#610 ·
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
gitbutlerapp/gitbutler#15998 · 1 comment ·
-
bug triage:deciding
Difficulty 1/5 Under an hour Newbie friendliness 88/100
open-telemetry/otel-arrow#4132 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100