Investigate possible missed admission-resumption wakeup with libuv before 1.53
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 35/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- cpp
- Domain
- backend, distributed-systems, performance
Research direction
Examine the ordering in src/tls/openssl_server.h around lines 1830-1843 and 1303-1316. Model the interaction between the atomic recheck_read_interest flag and libuv's pending flag. Validate the potential missed-notification sequence against libuv's documentation and implementation, particularly versions before 1.53. Determine if a fence or mutex protocol is needed to ensure correct ordering.
Written by the indexing model from the issue text.
Description
Concern
Investigate whether inbound admission resumption can miss a wakeup when libuv coalesces notifications. This is a portable memory-ordering concern found during review, not a reproduced failure.
The path exists on upstream 00095f2519a119ac3fd5c384c9064a9302ffc999, before the shared notification locking proposed in #8435.
Relevant ordering
When inbound capacity becomes available, the registered waker does:
recheck_read_interest.store(true, std::memory_order_release);
wake();
The host callback consumes that flag before acquiring out_mutex:
const bool recheck_all =
recheck_read_interest.exchange(false, std::memory_order_acq_rel);
Unlike pending responses and completions, publication and consumption of this flag do not share a mutex. Its atomic operations avoid a data race, but do not by themselves establish ordering with libuv's separate pending flag.
Potential missed-notification sequence
- Libuv clears its pending flag and starts the callback.
- The callback consumes
recheck_read_interestas false. - A producer sets
recheck_read_interestto true and callswake(). - If the older libuv coalescing path observes a stale pending value of one, it returns without scheduling another callback.
- The current callback finishes without rechecking admission.
The producer can take and release the lifecycle mutex before the callback's later lifecycle check. That check therefore does not establish the ordering needed before the producer's pending-flag load, nor does it recheck admission.
This sequence needs validation against the complete implementation and memory model. It has not been demonstrated on the deployed compiler, pthread implementation or CPU.
If reachable, paused socket reads could remain paused until another event causes an admission recheck. This would be a liveness issue, not lost queued response data. Its duration and practical reachability are unknown.
libuv contract
Libuv 1.48 documents thread-safe sends and notification coalescing, but not the stronger publication guarantee.
Current libuv documentation states that send/callback sequential consistency was added in 1.53.0 and warns that earlier coalescing cases may require a full sequentially consistent fence. The 1.48 implementation has a relaxed pending-flag fast-path load.
Investigation and acceptance
- Validate or reject the ordering concern with a reduced model covering the admission flag, both lifecycle-lock acquisitions and libuv's pending flag.
- Add a regression for admission recovery when unrelated traffic stops, including recovery triggered by a different interface.
- If confirmed, establish explicit publication/notification ordering for supported libuv versions. Possible approaches include a shared mutex protocol for this flag or the documented fence mitigation; verify the complete protocol before choosing.
- Keep this separate from #8435: the same concern applies to the existing exclusive notifier lock, and does not invalidate the mutex-protected response/completion queue protocol.
- Dominant language
- C++
- Stars
- 876
- Forks
- 260
- Avg merge
- 1d 13h
- Merged PRs (30d)
- 163
Getting set up
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/CCF
-
Difficulty 4/5 3-5 days Newbie friendliness 38/100
Maintainers usually reply within 1 day
-
Difficulty 3/5 1-2 days Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 45/100
microsoft/CCF#8439 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
Maintainers usually reply within 1 day
-
Difficulty 5/5 Over a week Newbie friendliness 30/100
microsoft/CCF#8184 · 1 comment ·
Maintainers usually reply within 1 day
Similar issues
-
Difficulty 1/5 Under an hour Newbie friendliness 92/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
cp-algorithms/cp-algorithms#1715 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Icinga/icinga2#11058 · 1 comment ·
Maintainers usually reply within 1 day
-
status:needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
PX4/PX4-Autopilot#28924 ·
Maintainers usually reply within 1 day
-
component: split-view platform: windows
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
zen-browser/desktop#15616 · 1 reaction ·
Maintainers usually reply within 1 day