Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Healthy Orleans silo falsely voted dead by peers and terminated itself during cluster roll-churn

Closed
#6,396 1 comment 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
8/100
Issue type
Bug
Clarity
Needs clarification
Activity status
Active
Tech stack
csharp, kubernetes

Research direction

Start with /_/src/Orleans.Runtime/MembershipService/MembershipTableManager.cs, at CheckIfLocalSiloIsDead and KillMyselfLocally, then the suspicion vote path in ClusterHealthMonitor. The report says the precise cause cannot be determined from logs and needs pod resource and GC metrics for the implicated pods. Done would need a reproduction or metrics that explain the missed probes, plus a decision on whether the fix is code or ClusterMembershipOptions tuning.

Written by the indexing model from the issue text.

Description

bug sev:M

What is failing

A live, actively heartbeating Orleans silo in the memex portal cluster was marked Dead in the membership table by two peer silos and honored the verdict by terminating itself (MembershipTableManager.KillMyselfLocally → FatalErrorHandler). The silo's last IAmAlive update postdates all three suspicion votes, so it was alive and writing to the membership table at the moment it was voted dead — a false-positive death declaration, not a genuine crash or rolling-shutdown race (the pod was not terminating; the silo had started minutes earlier).

Probable cause — medium confidence

The suspicion protocol (default: 2 distinct silo votes) fired because the target silo failed to answer direct probes from both peers within the freshness window, even though its table-written heartbeats continued. That asymmetry typically points to transient unreachability of the pod — CPU/threadpool starvation, a long GC pause, or a node-level network hiccup — rather than a dead process. The wider cluster was under heavy roll-churn at the same moment (see the SiloUnavailableException roll-churn bursts ending seconds before this event), so deployment churn / load is the likely trigger. The precise reason the probes were missed is not determinable from these logs — it needs pod resource metrics for the implicated pods around the event.

Impact

Single occurrence on a single pod at the time of this incident. One healthy portal pod killed itself; in-flight grain calls to it failed transiently (the companion SiloUnavailableException burst at the same moment), and the pod was restarted by Kubernetes. User impact was transient, but a membership protocol that kills healthy silos under load amplifies instability during deployments.

Where to look

  • Orleans.Runtime.MembershipService.MembershipTableManager — CheckIfLocalSiloIsDead / KillMyselfLocally (/_/src/Orleans.Runtime/MembershipService/MembershipTableManager.cs), and the peer-side suspicion path (ClusterHealthMonitor / vote accumulation) on the suspecting silos.
  • Pod resource and GC logs for the terminated pod and its two suspecting peers around the event window.
  • Related incidents: the FatalErrorHandler companion at Admin/_LogIncident/bfcac5a35d2571b4 (same event, carries the full stack), and the ongoing roll-churn family e.g. Admin/_LogIncident/roll-churn-1e98b289db877f91 — this false-positive kill may be another symptom of that churn rather than an independent defect.

If investigation shows this only occurs during roll-churn windows, it may be worth tuning ClusterMembershipOptions (probe timeouts / vote thresholds) for this deployment instead of chasing a code fix.


Evidence
Fingerprint 89dfb3b5f652b720
Category Orleans.Runtime.MembershipService.MembershipTableManager
Severity Error
Namespace memex
Pods memex-portal-deployment-5f6b9fc9b-hrdhq
Occurrences 1
First seen 2026-10-09 20:44:12Z
Last seen 2026-10-09 20:44:12Z
Routing not determined — no configured route matches the category Orleans.Runtime.MembershipService.MembershipTableManager. This repository is the configured fallback, not a finding about who owns the fault; the category names the LOGGER, which may not be the subject.
Recent log lines
2026-10-09 20:44:12Z memex-portal-deployment-5f6b9fc9b-hrdhq fail: Orleans.Runtime.MembershipService.MembershipTableManager[100627]
      I have been told I am dead, so this silo will stop! Reason: I should be Dead according to the membership table (in Refresh). Local entry: [SiloAddress=S10.244.17.110:11111:150582958 SiloName=Silo_9dc7e Status=Dead HostName=memex-portal-deployment-5f6b9fc9b-hrdhq ProxyPort=30000 RoleName= UpdateZone=0 FaultZone=0 StartTime=2026-10-09 20:35:58.118 GMT IAmAliveTime=2026-10-09 20:44:06.335 GMT Suspecters=[S10.244.3.225:11111:150580999, S10.244.2.180:11111:150580798, S10.244.3.225:11111:150580999] SuspectTimes=[2026-10-09 20:43:48.767 GMT, 2026-10-09 20:43:48.767 GMT, 2026-10-09 20:43:48.767 GMT]].

Opened automatically from Admin/_LogIncident/89dfb3b5f652b720. Recurrences are folded into this issue rather than opening new ones.
It also stands for the whole log site 3ce5feb1528fd58f: other fingerprints of this site fold in here as comments rather than opening tickets of their own.

Dominant language
C#
Stars
12
Forks
5
Avg merge
4h 15m
Merged PRs (30d)
969

Getting set up

  • No Dockerfile or Docker Compose file
  • Has a pull request template
  • No contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Systemorph/MeshWeaver

All issues in Systemorph/MeshWeaver

Similar issues

More C# issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.