Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

MessageHub disposal deadlock watchdog fires while a quiesce drain is still in flight, then the hub completes disposal normally

Open
#6,156 3 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
48/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
csharp
Domain
backend

Research direction

Start by locating the disposal-progress watchdog that emits log site MeshWeaver.Messaging.MessageHub[7317] in MeshWeaver.Messaging; the issue does not identify its source file or tests. Inspect how the no-progress window handles drainsInFlight > 0 and compare that with the normal drain completion path. Done means the watchdog no longer reports a deadlock for a drain that subsequently completes normally, while genuine deadlocks remain detectable.

Written by the indexing model from the issue text.

Description

bug sev:L

What is failing. The disposal watchdog in MeshWeaver.Messaging.MessageHub (log site event 7317) declared a "DISPOSAL DEADLOCK DETECTED" verdict on hub sync/h4lsOgLfWkSCSM-YmfmATg after 8 seconds of no teardown progress — while its own counters showed drainsInFlight=1, i.e. a quiesce drain was still legitimately in flight. The diagnostic line emitted immediately after the verdict shows the same hub at RunLevel=Dead, Disposal=Completed with an empty queue (drainsInFlight=0, draining=False): the drain finished on its own, so no deadlock existed at the moment the verdict fired.

Probable cause (medium confidence). The watchdog's no-progress window does not exclude the case where a drain is in flight and awaiting an external reply — its own text says the usual outstanding item is "a reply owed from outside this mesh" — and/or the 8-second window is shorter than the normal tail latency of such a wait. The completion line proves the stall self-resolved, which supports a false positive; what the drain was actually waiting on during those 8 seconds is not in the evidence, and the log's own wording concedes the verdict "does not name a cause". I could not locate the MessageHub source in the mesh, so the exact threshold and the in-flight-drain guard could not be read.

Impact. One occurrence, one pod (memex-portal-deployment-6cfcf78898-kss8q), a single instant on 2026-10-05, self-resolved in ~8 seconds; disposal was not forced and the hub reached Dead/Completed normally. No user-visible effect. The cost is in the watchdog's signal value: a fail-level "DEADLOCK DETECTED" that resolves itself invites an investigation every time it fires and trains readers to ignore the real ones.

Where to look. The disposal-progress watchdog behind log site MeshWeaver.Messaging.MessageHub[7317] in MeshWeaver.Messaging — specifically the condition that decides to emit the verdict while drainsInFlight > 0, and the no-progress window it measures against. Related design context on the teardown diagnostics and their earlier false-cause wording: A shutdown that stops halfway now says so instead of waiting. No existing issue in the mesh covers this symptom — not a duplicate of MeshWeaver#5988 (CI gate) or #5935 (PR review sweep).


Evidence
Fingerprint 1f9d72d6eb86189c
Category MeshWeaver.Messaging.MessageHub
Severity Error
Namespace memex
Pods memex-portal-deployment-6cfcf78898-kss8q
Occurrences 1
First seen 2026-10-05 15:58:40Z
Last seen 2026-10-05 15:58:40Z
Routing not determined — no configured route matches the category MeshWeaver.Messaging.MessageHub. This repository is the configured fallback, not a finding about who owns the fault; the category names the LOGGER, which may not be the subject.
Recent log lines
2026-10-05 15:58:40Z memex-portal-deployment-6cfcf78898-kss8q fail: MeshWeaver.Messaging.MessageHub[7317]
      DISPOSAL DEADLOCK DETECTED: Hub sync/h4lsOgLfWkSCSM-YmfmATg made no teardown progress for 00:00:08 (last progress: sync/h4lsOgLfWkSCSM-YmfmATg → Started). RunLevel=Started, queue depth 0, drainsInFlight=1, drainsAwaitingScheduler=0, draining=True, 0 turn(s) dequeued since Dispose(). No turn is on the block, the pump is not holding queued work, and this hub has no hosted hubs and no outstanding child-disposal join — so THIS VERDICT DOES NOT NAME A CAUSE. What is still outstanding is in the diagnostics below (pending callbacks are the usual one: a reply owed from outside this mesh). Disposal is NOT forced.
      Hub sync/h4lsOgLfWkSCSM-YmfmATg RunLevel=Dead Disposal=Completed Queue(buffer=0,deferred=0,drainsInFlight=0,openGates=0,draining=False,drainsAwaitingScheduler=0) Registrants=0
      

Opened automatically from Admin/_LogIncident/1f9d72d6eb86189c. Recurrences are folded into this issue rather than opening new ones.
It also stands for the whole log site a22da98fbbaeabc6: other fingerprints of this site fold in here as comments rather than opening tickets of their own.

Dominant language
C#
Stars
12
Forks
5
Avg merge
4h 1m
Merged PRs (30d)
979

Getting set up

  • No Dockerfile or Docker Compose file
  • Has a pull request template
  • No contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Systemorph/MeshWeaver

All issues in Systemorph/MeshWeaver

Similar issues

More C# issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.