Silo 5j9gx marked Dead by peers 26s after its last heartbeat and terminated itself — recurrence of the false-positive membership kill (#3053)
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 12/100
- Issue type
- Bug
- Clarity
- Needs clarification
- Activity status
- Active
- Tech stack
- csharp
- Domain
- distributed-systems
Research direction
Start in Orleans.Runtime/MembershipService/MembershipTableManager.cs at CheckIfLocalSiloIsDead and KillMyselfLocally, and trace both the Refresh and CleanupTableEntries call sites. Decide whether suspect votes older than the local IAmAliveTime should be discounted, which is a membership-protocol design question. The issue names no test and no concrete change, so done is not defined here. The work needs maintainer input on the intended behaviour before any code is written.
Written by the indexing model from the issue text.
Description
What is failing
An Orleans silo in memex-portal-deployment (Silo_b7daa, 10.244.6.128:11111:150624604, pod memex-portal-deployment-85587f8869-5j9gx) found itself Dead in the membership table during a Refresh pass and honoured the verdict by terminating itself (MembershipTableManager.KillMyselfLocally → FatalErrorHandler.OnFatalException; Exception: null and the Environment.get_StackTrace() frame confirm an intentional process stop, not a crash). The entry shows IAmAliveTime=08:29:36.513 against SuspectTimes=08:29:10.921 — the silo was still alive and heartbeating ~26 seconds after the votes that declared it dead. This is the same false-positive membership kill already triaged and filed as #3053 (and #3052) on 2026-09-02; this is a recurrence, not a new defect shape.
Probable cause — moderate-to-high confidence, same as #3053
The suspicion votes were cast from a stale/shared liveness snapshot, most likely during rolling-deploy churn: the pod's StartTime is 08:10:04 GMT, so the replica set was only ~20 minutes old when the kill fired — a deploy window. Two markers from the entry support the stale-suspect theory rather than a genuine crash:
- All three
SuspectTimesare identical to the millisecond (08:29:10.921), and one suspecter (10.244.17.58:11111:150624528) appears twice with the same timestamp — one snapshot, not three independent probe failures. (Identical pattern to the Sept 2 incident.) - The suspect votes predate the target's own live heartbeat by ~26 s; the target kept writing
IAmAliveand only exited at 08:30:03, ~53 s after being declared dead.
The open question carried over from #3053: the membership adoption path does not appear to discount suspect votes that are older than the target's own live IAmAliveTime. This time the verdict was consumed in Refresh (Sept 2 was CleanupTableEntries), so both call sites route through CheckIfLocalSiloIsDead → KillMyselfLocally with the same stale-suspect exposure. The alternative benign explanation — a long GC/CPU stall on 5j9gx making a genuine suspicion correct-but-harsh — cannot be excluded from this fingerprint alone; check the pod's own logs for GC pause warnings and the restart reason before ruling.
Impact
One occurrence, one pod, self-terminated and rescheduled by Kubernetes; the cluster remained quorate. However, this is the third firing of this shape (2026-09-02 ×2, now 2026-10-10) — it appears to recur around deploys, and each false kill cascades reactivation of every grain and pod-hub the silo hosted plus in-flight delivery failures. Watch the fingerprint at the next rolling deploy; at this rate it is a recurring deploy-time nuisance, not an outage.
Where to look
Orleans.Runtime.MembershipService.MembershipTableManager.CheckIfLocalSiloIsDead/KillMyselfLocally(/_/src/Orleans.Runtime/MembershipService/MembershipTableManager.cs) — verify whether suspect votes older than the target's liveIAmAliveTimeare (or should be) discounted beforeKillMyselfLocally; both theRefreshandCleanupTableEntriescall sites hit this.- Peer silo logs for 10.244.17.58 and 10.244.5.66 around 08:29:10 GMT — what made both suspect 5j9gx at the same instant during a ~20-minute-old replica set.
- Membership/probe timeout configuration for
memex-portal-deployment— the suspicion window appears too tight relative to silos that are demonstrably still heartbeating. - Prior triage of the same defect: incident
Admin/_LogIncident/727f38e432ffd92f(#3053) andAdmin/_LogIncident/34ce0ade2c08ca47(#3052).
Duplicate note: same component and symptom as #3053 / #3052 — this ticket documents a recurrence (new evidence, and a second call site: Refresh); a human may fold it into #3053 if that investigation is still open.
Evidence
| Fingerprint | 0803784e63467766 |
| Category | Orleans.Runtime.FatalErrorHandler |
| Severity | Error |
| Namespace | memex-cloud |
| Pods | memex-portal-deployment-85587f8869-5j9gx |
| Occurrences | 1 |
| First seen | 2026-10-10 08:30:03Z |
| Last seen | 2026-10-10 08:30:03Z |
| Routing | not determined — no configured route matches the category Orleans.Runtime.FatalErrorHandler. This repository is the configured fallback, not a finding about who owns the fault; the category names the LOGGER, which may not be the subject. |
Recent log lines
2026-10-10 08:30:03Z memex-portal-deployment-85587f8869-5j9gx fail: Orleans.Runtime.FatalErrorHandler[100002]
Fatal error from Orleans.Runtime.MembershipService.MembershipTableManager. Context: I have been told I am dead, so this silo will stop! Reason: I should be Dead according to the membership table (in Refresh). Local entry: [SiloAddress=S10.244.6.128:11111:150624604 SiloName=Silo_b7daa Status=Dead HostName=memex-portal-deployment-85587f8869-5j9gx ProxyPort=30000 RoleName= UpdateZone=0 FaultZone=0 StartTime=2026-10-10 08:10:04.851 GMT IAmAliveTime=2026-10-10 08:29:36.513 GMT Suspecters=[S10.244.17.58:11111:150624528, S10.244.5.66:11111:150623330, S10.244.17.58:11111:150624528] SuspectTimes=[2026-10-10 08:29:10.921 GMT, 2026-10-10 08:29:10.921 GMT, 2026-10-10 08:29:10.921 GMT]].
FATAL EXCEPTION from Orleans.Runtime.MembershipService.MembershipTableManager. Context: I have been told I am dead, so this silo will stop! Reason: I should be Dead according to the membership table (in Refresh). Local entry: [SiloAddress=S10.244.6.128:11111:150624604 SiloName=Silo_b7daa Status=Dead HostName=memex-portal-deployment-85587f8869-5j9gx ProxyPort=30000 RoleName= UpdateZone=0 FaultZone=0 StartTime=2026-10-10 08:10:04.851 GMT IAmAliveTime=2026-10-10 08:29:36.513 GMT Suspecters=[S10.244.17.58:11111:150624528, S10.244.5.66:11111:150623330, S10.244.17.58:11111:150624528] SuspectTimes=[2026-10-10 08:29:10.921 GMT, 2026-10-10 08:29:10.921 GMT, 2026-10-10 08:29:10.921 GMT]].. Exception: null.\nCurrent stack: at System.Environment.get_StackTrace()
at Orleans.Runtime.FatalErrorHandler.OnFatalException(Object sender, String context, Exception exception) in /_/src/Orleans.Runtime/Core/FatalErrorHandler.cs:line 32
at Orleans.Runtime.MembershipService.MembershipTableManager.KillMyselfLocally(String reason) in /_/src/Orleans.Runtime/MembershipService/MembershipTableManager.cs:line 708
at Orleans.Runtime.MembershipService.MembershipTableManager.CheckIfLocalSiloIsDead(String caller) in /_/src/Orleans.Runtime/MembershipService/Membershi…[truncated]
Opened automatically from Admin/_LogIncident/0803784e63467766. Recurrences are folded into this issue rather than opening new ones.
It also stands for the whole log site dfcbe511c74e943b: other fingerprints of this site fold in here as comments rather than opening tickets of their own.
- Dominant language
- C#
- Stars
- 12
- Forks
- 5
- Avg merge
- 4h 15m
- Merged PRs (30d)
- 969
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- No contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Systemorph/MeshWeaver
-
ApiTokenService.RevokeToken posts its revocation SaveMeshNodeRequest from the mesh (router) hub instead of a node-operation hubPossibly taken A pull request linked to this issue is open or already merged. Openbug sev:M
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Systemorph/MeshWeaver#6026 · 4 comments ·
Maintainers usually reply within 1 day
-
bug sev:L
Difficulty 1/5 1-3 hours Newbie friendliness 76/100
Systemorph/MeshWeaver#6011 · 3 comments ·
Maintainers usually reply within 1 day
-
bug sev:L
Difficulty 4/5 3-5 days Newbie friendliness 45/100
Systemorph/MeshWeaver#6450 ·
Maintainers usually reply within 1 day
-
bug sev:M
Difficulty 4/5 3-5 days Newbie friendliness 22/100
Systemorph/MeshWeaver#6429 · 2 comments ·
Maintainers usually reply within 1 day
-
bug issue-text-names-sentinel sev:M
Difficulty 4/5 3-5 days Newbie friendliness 6/100
Systemorph/MeshWeaver#6408 ·
Maintainers usually reply within 1 day
All issues in Systemorph/MeshWeaver
Similar issues
-
Bug: Hidden loading rings keep the compositor running, so Files uses 7-94% of a CPU core while idleOpen
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
files-community/Files#19026 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
BHoM/Civil3D_Toolkit#118 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
rjmurillo/moq.analyzers#1468 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 83/100
-
area:frontend FE P2
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
klasolsson81/jobbliggaren#2139 ·
Maintainers usually reply within 1 day