Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

CopyNode's detached reply chain outlives the 60 s RequestTimeout on the node-ops hub — a move fails (or falsely fails) with no recorded verdict

Open
#6,105 7 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
45/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
csharp
Domain
backend

Research direction

Start with the CopyNodeRequest handler on NodeOperationExecutionHub and the linked SelfAddressedRequests, Write Verdict Totality, and Copy Completeness architecture docs. Add hub.NoteRequestStage at the detached copy chain’s success, fault, and cancellation terminals, then verify each terminal is recorded; the issue does not name a test or test command.

Written by the indexing model from the issue text.

Description

bug sev:M

What is failing

A Move in the FeedbackHandover flow (a feedback submission being moved into Feedback/_Submissions) failed its 60-second RequestTimeout waiting for the CopyNodeRequest that a move issues on the mesh's single node-CRUD hub. The request is self-addressed — issuer and target are the same portal/nodeops hub — and the fate trail records every stage through HANDLER_EXIT state=Processed, then nothing for the full minute: a handler took the request, and the detached observable chain that owes the reply produced no reply, completion or fault.

Probable cause

The canonical node handlers return Processed() at once and answer later from a detached observable; for CopyNodeRequest that observable is the subtree copy itself. Two candidates remain, and today's instrumentation cannot split them:

  1. The copy ran past the bound. The hub was demonstrably alive during the wait (queue empty, RunLevel=Started, a non-zero handledWhileWaiting), so work still in flight at the 60 s caller-side deadline is plausible for a large subtree.
  2. The chain terminated without a terminal. The detached copy faulted or completed inside code that records no RequestFateLedger stage, so nobody answered and nobody ever would have.

Confidence: high that the reply is owed by the detached copy chain (intake, gate, routing and handler entry/exit are all recorded as succeeded); low on which of the two candidates applies until the missing stage instrumentation exists — the trail's own verdict asks for hub.NoteRequestStage(...) at the CopyNode handler's terminal arms.

Either way the consequence is the dangerous half of "a timed-out delivery is still held by the callee": the handover reports failure while the copy may still land. The sibling Node already exists incidents from the same FeedbackHandover flow in the same window look exactly like retries meeting their own earlier copy — treat a retry-on-timeout here as unsafe until the callee's verdict is observable.

Impact

One failed (or falsely-failed) feedback submission move on a single portal pod. User-visible failure of a secondary path with an easy workaround; the hub kept serving other traffic throughout the window, so this is not a saturation/wedge of the node-CRUD hub (the empty queue and live handledWhileWaiting rule that out).

Where to start


Evidence
Fingerprint a10ae477c7ada9c6
Category MeshWeaver.Mesh.Services.IMeshCatalog
Severity Error
Exception System.TimeoutException
Namespace memex
Pods memex-portal-deployment-76dd69c7fc-xzsdl
Occurrences 1
First seen 2026-10-04 20:15:03Z
Last seen 2026-10-04 20:15:03Z
Routing not determined — no configured route matches the category MeshWeaver.Mesh.Services.IMeshCatalog. This repository is the configured fallback, not a finding about who owns the fault; the category names the LOGGER, which may not be the subject.
Recent log lines
2026-10-04 20:15:03Z memex-portal-deployment-76dd69c7fc-xzsdl fail: MeshWeaver.Mesh.Services.IMeshCatalog[0]
      Move rbuergi/Feedback/fb-agentround-instanceaction-compile-20261004 -> Feedback/_Submissions/fb-feb34a3557f5b3e0d104193110375e19 failed
      System.TimeoutException: No response received in hub portal/nodeops--PztfvZ_ckWW2sZqk6eiZg within 00:01:00 for request CopyNodeRequest (id=MzAbFzyZJ0qmqstzp_n4iw) → target portal/nodeops--PztfvZ_ckWW2sZqk6eiZg. This hub: RunLevel=Started Queue(buffer=0,deferred=0,openGates=0,drainsInFlight=0,draining=False,handledWhileWaiting=78). 🚨 THIS HUB IS ALSO THE TARGET, so the request never left it: there is no routing leg that could have lost it and no reply leg that could have lost the answer, and "the target's own RunLevel and queue" are the numbers printed above. An empty queue here does NOT mean the request was never handled — the canonical mesh handlers return Processed() at once and owe their reply from a DETACHED observable, so a handler that ran and has not yet produced a terminal looks exactly like one that never ran. What is left is: the delivery was refused at this hub's own intake, or a handler took it and the work that owes the reply produced no terminal. The trail below says which. Trail: AWAITING CopyNodeRequest→portal/nodeops--PztfvZ_ckWW2sZqk6eiZg@portal/nodeops--PztfvZ_ckWW2sZqk6eiZg(+0ms) → POSTED target=portal/nodeops--PztfvZ_ckWW2sZqk6eiZg@portal/nodeops--PztfvZ_ckWW2sZqk6eiZg(+0ms) → RECEIVED runLevel=Started@portal/nodeops--PztfvZ_ckWW2sZqk6eiZg(+0ms) → ENQUEUED@portal/nodeops--PztfvZ_ckWW2sZqk6eiZg(+0ms) → QUEUED queue=main depth=1@portal/nodeops--PztfvZ_ckWW2sZqk6eiZg(+0ms) → ROUTED onTarget=True state=Submitted@portal/nodeops--PztfvZ_ckWW2sZqk6eiZg(+0ms) → HANDLER_ENTER@portal/nodeops--PztfvZ_ckWW2sZqk6eiZg(+0ms) → HANDLER_EXIT state=Processed@portal/nodeops--PztfvZ_ckWW2sZqk6eiZg(+0ms)  ⇒ a handler was entered and no reply, completion or fault has been recorded since — the work that owes the reply is either STILL RUNNING or terminated inside code that rec…[truncated]

Opened automatically from Admin/_LogIncident/a10ae477c7ada9c6. Recurrences are folded into this issue rather than opening new ones.
It also stands for the whole log site d834398a8b064677: other fingerprints of this site fold in here as comments rather than opening tickets of their own.

Dominant language
C#
Stars
12
Forks
5
Avg merge
3h 53m
Merged PRs (30d)
968

Getting set up

  • No Dockerfile or Docker Compose file
  • Has a pull request template
  • No contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Systemorph/MeshWeaver

All issues in Systemorph/MeshWeaver

Similar issues

More C# issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.