SessionPool acquire: growth/hand-out failures bypass acquire_timeout budget; live counter decremented outside the state lock loses Condvar wakeups; acquire_timeout = Duration::MAX panics

Open
#6 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
62/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
rust
Domain
databases

Research direction

Start by locating SessionPool::acquire and tracing the open_session growth path and hand_out USE replay path, then inspect every live mutation relative to the state lock and Condvar notification. Verify that failures retry until the original deadline, live decrements notify waiters under the lock, and Duration::MAX waits without overflow; reproduce the behavior with fake-listener stress tests.

Written by the indexing model from the issue text.

Description

bug

Three related pool-accounting defects found by stress-testing SessionPool against fake listeners:

  1. acquire() spends none of its acquire_timeout budget when the growth branch fails. When open_session() fails, acquire() returns the error immediately (measured 198us against a 5s budget; 4136 of 4800 acquires failed instantly under stress while sessions were circulating). The same happens when the USE replay in hand_out fails — it discards the remaining idle candidates instead of retrying the loop.
  2. live is decremented outside the state lock at several sites, and the decremented value is exactly the predicate parked waiters re-evaluate, so Condvar notifications can be lost. The comment "Only mutated while holding state" no longer matches the code. Measured stalls track acquire_timeout exactly (300ms → 305ms, 1000ms → 1.005s).
  3. acquire_timeout = Duration::MAX panics on Instant::now() + acquire_timeout overflow before the lock is taken. There is no "never time out" option, and Duration::MAX is the natural way to ask for one.

Fix: on growth/hand-out failure, re-take the lock, decrement live under it, notify, and continue the loop (bounded by the original deadline); use a saturating deadline (checked_add) so Duration::MAX means "wait without timeout".

Dominant language
Rust
Stars
1
Forks
0
PR merge metrics
No merged PRs in 30d

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from apache/iotdb-client-rust

All issues in apache/iotdb-client-rust

Similar issues

More Rust issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.