Healthy Orleans silo falsely voted dead by peers and terminated itself during cluster roll-churn
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 8/100
- Issue type
- Bug
- Clarity
- Needs clarification
- Activity status
- Active
- Tech stack
- csharp, kubernetes
- Domain
- distributed-systems
Research direction
Start with /_/src/Orleans.Runtime/MembershipService/MembershipTableManager.cs, at CheckIfLocalSiloIsDead and KillMyselfLocally, then the suspicion vote path in ClusterHealthMonitor. The report says the precise cause cannot be determined from logs and needs pod resource and GC metrics for the implicated pods. Done would need a reproduction or metrics that explain the missed probes, plus a decision on whether the fix is code or ClusterMembershipOptions tuning.
Written by the indexing model from the issue text.
Description
What is failing
A live, actively heartbeating Orleans silo in the memex portal cluster was marked Dead in the membership table by two peer silos and honored the verdict by terminating itself (MembershipTableManager.KillMyselfLocally → FatalErrorHandler). The silo's last IAmAlive update postdates all three suspicion votes, so it was alive and writing to the membership table at the moment it was voted dead — a false-positive death declaration, not a genuine crash or rolling-shutdown race (the pod was not terminating; the silo had started minutes earlier).
Probable cause — medium confidence
The suspicion protocol (default: 2 distinct silo votes) fired because the target silo failed to answer direct probes from both peers within the freshness window, even though its table-written heartbeats continued. That asymmetry typically points to transient unreachability of the pod — CPU/threadpool starvation, a long GC pause, or a node-level network hiccup — rather than a dead process. The wider cluster was under heavy roll-churn at the same moment (see the SiloUnavailableException roll-churn bursts ending seconds before this event), so deployment churn / load is the likely trigger. The precise reason the probes were missed is not determinable from these logs — it needs pod resource metrics for the implicated pods around the event.
Impact
Single occurrence on a single pod at the time of this incident. One healthy portal pod killed itself; in-flight grain calls to it failed transiently (the companion SiloUnavailableException burst at the same moment), and the pod was restarted by Kubernetes. User impact was transient, but a membership protocol that kills healthy silos under load amplifies instability during deployments.
Where to look
Orleans.Runtime.MembershipService.MembershipTableManager—CheckIfLocalSiloIsDead/KillMyselfLocally(/_/src/Orleans.Runtime/MembershipService/MembershipTableManager.cs), and the peer-side suspicion path (ClusterHealthMonitor/ vote accumulation) on the suspecting silos.- Pod resource and GC logs for the terminated pod and its two suspecting peers around the event window.
- Related incidents: the
FatalErrorHandlercompanion atAdmin/_LogIncident/bfcac5a35d2571b4(same event, carries the full stack), and the ongoing roll-churn family e.g.Admin/_LogIncident/roll-churn-1e98b289db877f91— this false-positive kill may be another symptom of that churn rather than an independent defect.
If investigation shows this only occurs during roll-churn windows, it may be worth tuning ClusterMembershipOptions (probe timeouts / vote thresholds) for this deployment instead of chasing a code fix.
Evidence
| Fingerprint | 89dfb3b5f652b720 |
| Category | Orleans.Runtime.MembershipService.MembershipTableManager |
| Severity | Error |
| Namespace | memex |
| Pods | memex-portal-deployment-5f6b9fc9b-hrdhq |
| Occurrences | 1 |
| First seen | 2026-10-09 20:44:12Z |
| Last seen | 2026-10-09 20:44:12Z |
| Routing | not determined — no configured route matches the category Orleans.Runtime.MembershipService.MembershipTableManager. This repository is the configured fallback, not a finding about who owns the fault; the category names the LOGGER, which may not be the subject. |
Recent log lines
2026-10-09 20:44:12Z memex-portal-deployment-5f6b9fc9b-hrdhq fail: Orleans.Runtime.MembershipService.MembershipTableManager[100627]
I have been told I am dead, so this silo will stop! Reason: I should be Dead according to the membership table (in Refresh). Local entry: [SiloAddress=S10.244.17.110:11111:150582958 SiloName=Silo_9dc7e Status=Dead HostName=memex-portal-deployment-5f6b9fc9b-hrdhq ProxyPort=30000 RoleName= UpdateZone=0 FaultZone=0 StartTime=2026-10-09 20:35:58.118 GMT IAmAliveTime=2026-10-09 20:44:06.335 GMT Suspecters=[S10.244.3.225:11111:150580999, S10.244.2.180:11111:150580798, S10.244.3.225:11111:150580999] SuspectTimes=[2026-10-09 20:43:48.767 GMT, 2026-10-09 20:43:48.767 GMT, 2026-10-09 20:43:48.767 GMT]].
Opened automatically from Admin/_LogIncident/89dfb3b5f652b720. Recurrences are folded into this issue rather than opening new ones.
It also stands for the whole log site 3ce5feb1528fd58f: other fingerprints of this site fold in here as comments rather than opening tickets of their own.
- Dominant language
- C#
- Stars
- 12
- Forks
- 5
- Avg merge
- 4h 15m
- Merged PRs (30d)
- 969
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- No contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Systemorph/MeshWeaver
-
ApiTokenService.RevokeToken posts its revocation SaveMeshNodeRequest from the mesh (router) hub instead of a node-operation hubPossibly taken A pull request linked to this issue is open or already merged. Openbug sev:M
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Systemorph/MeshWeaver#6026 · 4 comments ·
Maintainers usually reply within 1 day
-
bug sev:L
Difficulty 1/5 1-3 hours Newbie friendliness 76/100
Systemorph/MeshWeaver#6011 · 3 comments ·
Maintainers usually reply within 1 day
-
bug sev:L
Difficulty 4/5 3-5 days Newbie friendliness 45/100
Systemorph/MeshWeaver#6450 ·
Maintainers usually reply within 1 day
-
bug sev:M
Difficulty 5/5 Over a week Newbie friendliness 12/100
Systemorph/MeshWeaver#6432 · 4 comments ·
Maintainers usually reply within 1 day
-
bug sev:M
Difficulty 4/5 3-5 days Newbie friendliness 22/100
Systemorph/MeshWeaver#6429 · 2 comments ·
Maintainers usually reply within 1 day
All issues in Systemorph/MeshWeaver
Similar issues
-
[誤判定] `define` が `デフィね`・`デフィ値` になるPossibly taken A pull request linked to this issue is open or already merged. Open再現済み 要トリアージ 誤判定
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
yksr-melt/Meltype#421 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 64/100
Facepunch/sbox-public#12063 · 1 comment ·
Maintainers usually reply within 2 days
-
documentation
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
facioquo/stock-indicators-dotnet#2316 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
lostindark/DriverStoreExplorer#477 ·
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
eriknihlen/OpenAC#219 ·
Maintainers usually reply within 1 day