OrleansRoutingService.DeliverMessage OOMs while routing a delivery — application-frame stack, distinct from the DrainLoop OOM chain (#4824/#5555/#6161)
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 48/100
Research direction
Start with src/MeshWeaver.Connection.Orleans/OrleansRoutingService.cs, especially the DeliverMessage closure at the reported lines 523 and 562, and identify the ImmutableHashSet being unioned. Compare hub routing-set growth across reconnects with pod memory telemetry; read src/MeshWeaver.Messaging.Hub/MessageService.cs at DrainLoop() and NotifyAsync() for delivery containment. Done when evidence distinguishes process-wide heap exhaustion from unbounded routing-set growth and the resulting failure or containment work is verified.
Written by the indexing model from the issue text.
Description
What is failing
While delivering a message to a portal client's mesh hub (mesh/{guid}), the MeshWeaver.Messaging.MessageService delivery pipeline aborted with System.OutOfMemoryException inside MeshWeaver.Connection.Orleans's OrleansRoutingService.DeliverMessage reactive step (RouteMessageAsync → Rx Defer → the DeliverMessage closure). The affected delivery is lost and not retried.
Probable cause
Medium confidence, two candidate layers that the logs alone cannot separate:
- Process-wide heap exhaustion (the drain loop as victim). The second occurrence (2026-10-05 16:33:18Z, pod
…gkkdv) landed in a five-minute window in which the same pod also OOM'd at two unrelated sites — aPrBabysitterdeserialization OOM at 16:28:03Z (filed as #6158) and aDrainLoop-stack delivery OOM at 16:32:38Z (sibling incidentf6a4dfaef500b486, filed as #6161). Allocation failing across three unrelated components in minutes is the signature of a pod at its memory ceiling, not of one bad allocation. - Unbounded growth of the hub's routing set. Unlike the #4824/#5555/#6161 chain — whose stacks ran entirely through
System.Reactivewith no application frame — this burst's top frame is application code, and the first sample shows the failing allocation insideImmutableHashSet<T>.Union(OrleansRoutingService.cs:562; the second sample fails at line 523 of the same closure). That is consistent with the set being unioned growing without bound (e.g. accumulating re-registrations) until the allocation fails — which would make this a distinct, routing-side defect rather than a pure heap victim.
Pod memory telemetry around 2026-10-05 16:28–16:33Z and the size of the affected hub's routing set at failure would discriminate between the two.
Impact
Two occurrences over ~21 hours (2026-10-04 19:45Z → 2026-10-05 16:33Z) across two memex portal pods. Intermittent, each costing one lost delivery — but the pod that failed last (gkkdv) was failing allocations across at least three components within five minutes, i.e. a degrading pod until restart. sev:H on the same reasoning as #6161: intermittent user-visible delivery failure, and evidence of silent pod degradation; a human may re-rule.
Where to look
src/MeshWeaver.Connection.Orleans/OrleansRoutingService.cs— theDeliverMessageclosure (lines 523 and 562 in the two samples): whichImmutableHashSetis beingUnioned, and whether registrations on a hub's routing set can accumulate without bound across client reconnects.src/MeshWeaver.Messaging.Hub/MessageService.cs—DrainLoop()/NotifyAsync()(lines 1713/1886 in the stacks) for the per-delivery containment question; the containment fix for the sibling recurrence is still parked (see the round-17 note on the #5555 triage item).- Related, not duplicates: #6161 (same category and message, pure-Rx stack — the earlier chain) and #6158 (sibling OOM on the same pod). If the heap-exhaustion hypothesis holds, all three share one root cause and can be folded into a single investigation; the routing-set growth candidate applies only to this incident.
Evidence
| Fingerprint | b696063f772f8b94 |
| Category | MeshWeaver.Messaging.MessageService |
| Severity | Error |
| Exception | System.OutOfMemoryException |
| Top frame | MeshWeaver.Connection.Orleans.OrleansRoutingService.<>c__DisplayClass27_0.<DeliverMessage>b__0() |
| Namespace | memex |
| Pods | memex-portal-deployment-76dd69c7fc-xxql9, memex-portal-deployment-6cfcf78898-gkkdv |
| Occurrences | 2 |
| First seen | 2026-10-04 19:45:08Z |
| Last seen | 2026-10-05 16:33:18Z |
| Routing | not determined — no configured route matches the category MeshWeaver.Messaging.MessageService. This repository is the configured fallback, not a finding about who owns the fault; the category names the LOGGER, which may not be the subject. |
Recent log lines
2026-10-04 19:45:08Z memex-portal-deployment-76dd69c7fc-xxql9 fail: MeshWeaver.Messaging.MessageService[0]
Unhandled exception in delivery pipeline for hub mesh/b6_J2gWCKkOSQzyRtKMAbA
System.OutOfMemoryException: Exception of type 'System.OutOfMemoryException' was thrown.
at System.Collections.Immutable.ImmutableHashSet`1.Union(IEnumerable`1 other, MutationInput origin)
at MeshWeaver.Connection.Orleans.OrleansRoutingService.<>c__DisplayClass27_0.<DeliverMessage>b__0() in /home/runner/work/MeshWeaver/MeshWeaver/src/MeshWeaver.Connection.Orleans/OrleansRoutingService.cs:line 562
at System.Reactive.Linq.ObservableImpl.Defer`1._.Run()
--- End of stack trace from previous location ---
at System.Reactive.PlatformServices.ExceptionServicesImpl.Rethrow(Exception exception)
at System.Reactive.ExceptionHelpers.Throw(Exception exception)
at System.Reactive.Stubs.<>c.<.cctor>b__2_1(Exception ex)
at System.Reactive.AnonymousSafeObserver`1.OnError(Exception error)
at System.Reactive.Sink`1.ForwardOnError(Exception error)
at System.Reactive.Linq.ObservableImpl.Defer`1._.Run()
at System.Reactive.Concurrency.CurrentThreadScheduler.Schedule[TState](TState state, TimeSpan dueTime, Func`3 action)
at System.Reactive.Concurrency.Scheduler.ScheduleAction[TState](IScheduler scheduler, TState state, Action`1 action)
at System.Reactive.Producer`2.SubscribeRaw(IObserver`1 observer, Boolean enableSafeguard)
at MeshWeaver.Messaging.HierarchicalRouting.RouteMessageAsync(IMessageDelivery delivery, CancellationToken cancellationToken)
at MeshWeaver.Messaging.MessageService.NotifyAsync(IMessageDelivery delivery, CancellationToken cancellationToken, Int64 turnSeq) in /home/runner/work/MeshWeaver/MeshWeaver/src/MeshWeaver.Messaging.Hub/MessageService.cs:line 1886
at MeshWeaver.Messaging.MessageService.DrainLoop() in /home/runner/work/MeshWeaver/MeshWeaver/src/MeshWeaver.Messaging.Hub/MessageService.cs:line 1713
2026-10-05 16:33:18Z memex-portal-deployment-6cfcf78898-gkkdv fail: MeshWeaver.Messaging.MessageService[0]
Unhandled exception in delivery pipeline for hub mesh/znseJ1iFZUChd0XTK-lwNA
System.OutOfMemoryException: Exception of type 'System.OutOfMemoryException' was thrown.
at MeshWeaver.Connection.Orleans.OrleansRoutingService.<>c__DisplayClass27_0.<DeliverMessage>b__0() in /home/runner/work/MeshWeaver/MeshWeaver/src/MeshWeaver.Connection.Orleans/OrleansRoutingService.cs:line 523
at System.Reactive.Linq.ObservableImpl.Defer`1._.Run()
--- End of stack trace from previous location ---
at System.Reactive.PlatformServices.ExceptionServicesImpl.Rethrow(Exception exception)
at System.Reactive.ExceptionHelpers.Throw(Exception exception)
at System.Reactive.Stubs.<>c.<.cctor>b__2_1(Exception ex)
at System.Reactive.AnonymousSafeObserver`1.OnError(Exception error)
at System.Reactive.Sink`1.ForwardOnError(Exception error)
at System.Reactive.Linq.ObservableImpl.Defer`1._.Run()
at System.Reactive.Concurrency.CurrentThreadScheduler.Schedule[TState](TState state, TimeSpan dueTime, Func`3 action)
at System.Reactive.Concurrency.Scheduler.ScheduleAction[TState](IScheduler scheduler, TState state, Action`1 action)
at System.Reactive.Producer`2.SubscribeRaw(IObserver`1 observer, Boolean enableSafeguard)
at MeshWeaver.Messaging.HierarchicalRouting.RouteMessageAsync(IMessageDelivery delivery, CancellationToken cancellationToken)
at MeshWeaver.Messaging.MessageService.NotifyAsync(IMessageDelivery delivery, CancellationToken cancellationToken, Int64 turnSeq) in /home/runner/work/MeshWeaver/MeshWeaver/src/MeshWeaver.Messaging.Hub/MessageService.cs:line 1886
at MeshWeaver.Messaging.MessageService.DrainLoop() in /home/runner/work/MeshWeaver/MeshWeaver/src/MeshWeaver.Messaging.Hub/MessageService.cs:line 1713
Opened automatically from Admin/_LogIncident/b696063f772f8b94. Recurrences are folded into this issue rather than opening new ones.
It also stands for the whole log site 0bd62bc50920f9e0: other fingerprints of this site fold in here as comments rather than opening tickets of their own.
- Dominant language
- C#
- Stars
- 12
- Forks
- 5
- Avg merge
- 3h 53m
- Merged PRs (30d)
- 968
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- No contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Systemorph/MeshWeaver
-
sev:L
Difficulty 1/5 Under an hour Newbie friendliness 72/100
Systemorph/MeshWeaver#6233 ·
Maintainers usually reply within 1 day
-
documentation feedback sev:L
Difficulty 1/5 Under an hour Newbie friendliness 82/100
Systemorph/MeshWeaver#6033 ·
Maintainers usually reply within 1 day
-
area:search documentation
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Systemorph/MeshWeaver#6030 ·
Maintainers usually reply within 1 day
-
bug sev:M
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Systemorph/MeshWeaver#6026 · 1 comment ·
Maintainers usually reply within 1 day
-
area:hosting bug sev:L
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Systemorph/MeshWeaver#6025 ·
Maintainers usually reply within 1 day
All issues in Systemorph/MeshWeaver
Similar issues
-
Bug pulumi/pulumi
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
activescott/lessmsi#306 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Maintainers usually reply within 1 day
-
HTML sitemap lists unpublished pagesPossibly taken @KrzysztofPajak claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
grandnode/grandnode2#883 ·
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Maintainers usually reply within 1 day