Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Fleet coordinator cannot create its Hosting/Coordinator node — the runtime fault-loops every minute and no coordinator round has ever run

Closed
#6,150 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
38/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
csharp
Domain
backend

Research direction

Read Hosting/Coordination/Source/FleetCoordinatorRuntime.cs, focusing on EnsureNode and the AddFleetCoordinator SelfHealing wrapper; compare deployment and closed-type registration steps in the referenced issues #2871 and #2898. First determine whether the missing node reflects incomplete deployment or a swallowed create refusal. Done means the coordinator no longer faults every minute and create failures are surfaced distinctly with an appropriate recovery path.

Written by the indexing model from the issue text.

Description

bug sev:M

What is failing

The fleet coordinator (Plugins #2871, merged 2026-10-05 08:52Z) is armed on the control instance's always-on Hosting/PlatformBuilds hub, but it cannot bring up its own node Hosting/Coordinator. Every minute the schedule faults with MeshWeaver.Messaging.DeliveryFailureException — [ADDRESS] node found at 'Hosting/Coordinator'. Closest ancestor is 'Hosting'" — and restarts, so the coordinator has never completed a single round: no startup round, no fleet ranking, no dispatch, since it was deployed.

Probable cause (high confidence)

Hosting/Coordinator is created at run time with NodeType Hosting/Coordination, and that type is not usable on the running control instance:

  • The NodeType definition node does not exist in the mesh (no Hosting/Coordination anywhere under Hosting — every sibling type like Hosting/Queue, Hosting/RepoHealth is present), so [PERSON_NAME]Hostingthat was to carryHosting/Coordination.json` (#2871's deploy step 1) appears not to have landed.
  • Even with a sync, the control instance runs Mesh:ClosedTypeSet, and #2871's registration of Hosting/Coordination in src/MeshWeaver.Fleet.Control/Generated/FleetControlNodeTypes.g.cs only takes effect after the control image is rebuilt and Rolled — exactly the state PR #2898 documented for Hosting/PrWatch: "until then the running control pods keep the old registration set", and an unregistered type is refused at activation.

The aggravating code defect: FleetCoordinatorRuntime.EnsureNode wraps the create in .Catch(_ => Observable.Return(Unit.Default)) and swallows the refusal silently. The round loop then combines over GetMeshNodeStream("Hosting/Coordinator"), which faults, and the SelfHealing(..., TimeSpan.FromMinutes(1)) wrapper logs the generic "the schedule faulted… restarting in a minute" line — the one-per-minute cadence in this incident is that self-healing restart. Nothing in the logs distinguishes "the create was refused" from any other fault, so the runtime spins forever instead of surfacing the deployment gap.

Impact

The coordinator is a brand-new automation merged this morning; nothing that worked before is broken by its absence. What is affected: the coordinator's whole function (priority ranking and load-aware dispatch across the fleet's triage work) is dead on the control instance, and the control pod emits one error per minute indefinitely. Hosting:Coordinator:Enabled=false stops the noise if needed. Judged from the numbers this is a clear defect with an easy workaround, not an outage.

Where to look

  • Hosting/Triage/pull-request/systemorph-meshweaver-plugins-2871 — the coordinator's design and deploy steps: [PERSON_NAME] Hosting, recycle Hosting/PlatformBuildInbox and Hosting/Coordination. Note Hosting/Coordination is absent from the mesh while the runtime is clearly armed — the deploy is half-done.
  • Hosting/Triage/pull-request/systemorph-meshweaver-plugins-2898 — the same class of failure (an unregistered NodeType refused under the closed set) for Hosting/PrWatch, fixed by registering the type + image rebuild + Roll.
  • Hosting/Coordination/Source/FleetCoordinatorRuntime.cs — EnsureNode (the silent .Catch on the create) and AddFleetCoordinator's SelfHealing wrapper that turns the fault into the per-minute restart.

Suggested fix: complete the deploy (sync/register Hosting/Coordination, rebuild and roll the control image per #2898's pattern), and independently make EnsureNode log the create failure distinctly and back off instead of letting the loop fault-loop on a missing node — so a half-deployed coordinator says so instead of impersonating a schedule bug.


Evidence
Fingerprint 14487921ad43ddc7
Category FleetCoordinator
Severity Error
Exception MeshWeaver.Messaging.DeliveryFailureException
Namespace memex
Pods memex-portal-deployment-7c4764b686-tbr7f, memex-portal-deployment-58475df96-nprkt
Occurrences 17
First seen 2026-10-05 12:05:43Z
Last seen 2026-10-05 12:20:11Z
Routing not determined — no configured route matches the category FleetCoordinator. This repository is the configured fallback, not a finding about who owns the fault; the category names the LOGGER, which may not be the subject.
Recent log lines
2026-10-05 12:12:44Z memex-portal-deployment-7c4764b686-tbr7f fail: FleetCoordinator[0]
      [Coordinator] the schedule faulted on Hosting/PlatformBuilds — restarting in a minute
      MeshWeaver.Messaging.DeliveryFailureException: No node found at 'Hosting/Coordinator'. Closest ancestor is 'Hosting' (remainder='Coordinator').
2026-10-05 12:13:44Z memex-portal-deployment-7c4764b686-tbr7f fail: FleetCoordinator[0]
      [Coordinator] the schedule faulted on Hosting/PlatformBuilds — restarting in a minute
      MeshWeaver.Messaging.DeliveryFailureException: No node found at 'Hosting/Coordinator'. Closest ancestor is 'Hosting' (remainder='Coordinator').
2026-10-05 12:14:44Z memex-portal-deployment-7c4764b686-tbr7f fail: FleetCoordinator[0]
      [Coordinator] the schedule faulted on Hosting/PlatformBuilds — restarting in a minute
      MeshWeaver.Messaging.DeliveryFailureException: No node found at 'Hosting/Coordinator'. Closest ancestor is 'Hosting' (remainder='Coordinator').
2026-10-05 12:14:55Z memex-portal-deployment-58475df96-nprkt fail: FleetCoordinator[0]
      [Coordinator] the schedule faulted on content-type-registration/e147a6e18ee84db4a200198c414b4547 — restarting in a minute
      MeshWeaver.Messaging.DeliveryFailureException: No node found at 'Hosting/Coordinator'. Closest ancestor is 'Hosting' (remainder='Coordinator').
2026-10-05 12:15:11Z memex-portal-deployment-58475df96-nprkt fail: FleetCoordinator[0]
      [Coordinator] the schedule faulted on Hosting/PlatformBuilds — restarting in a minute
      MeshWeaver.Messaging.DeliveryFailureException: No node found at 'Hosting/Coordinator'. Closest ancestor is 'Hosting' (remainder='Coordinator').
2026-10-05 12:16:11Z memex-portal-deployment-58475df96-nprkt fail: FleetCoordinator[0]
      [Coordinator] the schedule faulted on Hosting/PlatformBuilds — restarting in a minute
      MeshWeaver.Messaging.DeliveryFailureException: No node found at 'Hosting/Coordinator'. Closest ancestor is 'Hosting' (remainder='Coordinator').
2026-10-05 12:17:11Z memex-portal-deployment-58475df96-nprkt fail: FleetCoordinator[0]
      [Coordinator] the schedule faulted on Hosting/PlatformBuilds — restarting in a minute
      MeshWeaver.Messaging.DeliveryFailureException: No node found at 'Hosting/Coordinator'. Closest ancestor is 'Hosting' (remainder='Coordinator').
2026-10-05 12:18:11Z memex-portal-deployment-58475df96-nprkt fail: FleetCoordinator[0]
      [Coordinator] the schedule faulted on Hosting/PlatformBuilds — restarting in a minute
      MeshWeaver.Messaging.DeliveryFailureException: No node found at 'Hosting/Coordinator'. Closest ancestor is 'Hosting' (remainder='Coordinator').
2026-10-05 12:19:11Z memex-portal-deployment-58475df96-nprkt fail: FleetCoordinator[0]
      [Coordinator] the schedule faulted on Hosting/PlatformBuilds — restarting in a minute
      MeshWeaver.Messaging.DeliveryFailureException: No node found at 'Hosting/Coordinator'. Closest ancestor is 'Hosting' (remainder='Coordinator').
2026-10-05 12:20:11Z memex-portal-deployment-58475df96-nprkt fail: FleetCoordinator[0]
      [Coordinator] the schedule faulted on Hosting/PlatformBuilds — restarting in a minute
      MeshWeaver.Messaging.DeliveryFailureException: No node found at 'Hosting/Coordinator'. Closest ancestor is 'Hosting' (remainder='Coordinator').

Opened automatically from Admin/_LogIncident/14487921ad43ddc7. Recurrences are folded into this issue rather than opening new ones.
It also stands for the whole log site facaac0ee8f27d84: other fingerprints of this site fold in here as comments rather than opening tickets of their own.

Dominant language
C#
Stars
12
Forks
5
Avg merge
4h 1m
Merged PRs (30d)
979

Getting set up

  • No Dockerfile or Docker Compose file
  • Has a pull request template
  • No contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Systemorph/MeshWeaver

All issues in Systemorph/MeshWeaver

Similar issues

More C# issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.