Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Silo 5j9gx marked Dead by peers 26s after its last heartbeat and terminated itself — recurrence of the false-positive membership kill (#3053)

Open
#6,432 4 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
12/100
Issue type
Bug
Clarity
Needs clarification
Activity status
Active
Tech stack
csharp

Research direction

Start in Orleans.Runtime/MembershipService/MembershipTableManager.cs at CheckIfLocalSiloIsDead and KillMyselfLocally, and trace both the Refresh and CleanupTableEntries call sites. Decide whether suspect votes older than the local IAmAliveTime should be discounted, which is a membership-protocol design question. The issue names no test and no concrete change, so done is not defined here. The work needs maintainer input on the intended behaviour before any code is written.

Written by the indexing model from the issue text.

Description

bug sev:M

What is failing

An Orleans silo in memex-portal-deployment (Silo_b7daa, 10.244.6.128:11111:150624604, pod memex-portal-deployment-85587f8869-5j9gx) found itself Dead in the membership table during a Refresh pass and honoured the verdict by terminating itself (MembershipTableManager.KillMyselfLocally → FatalErrorHandler.OnFatalException; Exception: null and the Environment.get_StackTrace() frame confirm an intentional process stop, not a crash). The entry shows IAmAliveTime=08:29:36.513 against SuspectTimes=08:29:10.921 — the silo was still alive and heartbeating ~26 seconds after the votes that declared it dead. This is the same false-positive membership kill already triaged and filed as #3053 (and #3052) on 2026-09-02; this is a recurrence, not a new defect shape.

Probable cause — moderate-to-high confidence, same as #3053

The suspicion votes were cast from a stale/shared liveness snapshot, most likely during rolling-deploy churn: the pod's StartTime is 08:10:04 GMT, so the replica set was only ~20 minutes old when the kill fired — a deploy window. Two markers from the entry support the stale-suspect theory rather than a genuine crash:

  • All three SuspectTimes are identical to the millisecond (08:29:10.921), and one suspecter (10.244.17.58:11111:150624528) appears twice with the same timestamp — one snapshot, not three independent probe failures. (Identical pattern to the Sept 2 incident.)
  • The suspect votes predate the target's own live heartbeat by ~26 s; the target kept writing IAmAlive and only exited at 08:30:03, ~53 s after being declared dead.

The open question carried over from #3053: the membership adoption path does not appear to discount suspect votes that are older than the target's own live IAmAliveTime. This time the verdict was consumed in Refresh (Sept 2 was CleanupTableEntries), so both call sites route through CheckIfLocalSiloIsDead → KillMyselfLocally with the same stale-suspect exposure. The alternative benign explanation — a long GC/CPU stall on 5j9gx making a genuine suspicion correct-but-harsh — cannot be excluded from this fingerprint alone; check the pod's own logs for GC pause warnings and the restart reason before ruling.

Impact

One occurrence, one pod, self-terminated and rescheduled by Kubernetes; the cluster remained quorate. However, this is the third firing of this shape (2026-09-02 ×2, now 2026-10-10) — it appears to recur around deploys, and each false kill cascades reactivation of every grain and pod-hub the silo hosted plus in-flight delivery failures. Watch the fingerprint at the next rolling deploy; at this rate it is a recurring deploy-time nuisance, not an outage.

Where to look

  • Orleans.Runtime.MembershipService.MembershipTableManager.CheckIfLocalSiloIsDead / KillMyselfLocally (/_/src/Orleans.Runtime/MembershipService/MembershipTableManager.cs) — verify whether suspect votes older than the target's live IAmAliveTime are (or should be) discounted before KillMyselfLocally; both the Refresh and CleanupTableEntries call sites hit this.
  • Peer silo logs for 10.244.17.58 and 10.244.5.66 around 08:29:10 GMT — what made both suspect 5j9gx at the same instant during a ~20-minute-old replica set.
  • Membership/probe timeout configuration for memex-portal-deployment — the suspicion window appears too tight relative to silos that are demonstrably still heartbeating.
  • Prior triage of the same defect: incident Admin/_LogIncident/727f38e432ffd92f (#3053) and Admin/_LogIncident/34ce0ade2c08ca47 (#3052).

Duplicate note: same component and symptom as #3053 / #3052 — this ticket documents a recurrence (new evidence, and a second call site: Refresh); a human may fold it into #3053 if that investigation is still open.


Evidence
Fingerprint 0803784e63467766
Category Orleans.Runtime.FatalErrorHandler
Severity Error
Namespace memex-cloud
Pods memex-portal-deployment-85587f8869-5j9gx
Occurrences 1
First seen 2026-10-10 08:30:03Z
Last seen 2026-10-10 08:30:03Z
Routing not determined — no configured route matches the category Orleans.Runtime.FatalErrorHandler. This repository is the configured fallback, not a finding about who owns the fault; the category names the LOGGER, which may not be the subject.
Recent log lines
2026-10-10 08:30:03Z memex-portal-deployment-85587f8869-5j9gx fail: Orleans.Runtime.FatalErrorHandler[100002]
      Fatal error from Orleans.Runtime.MembershipService.MembershipTableManager. Context: I have been told I am dead, so this silo will stop! Reason: I should be Dead according to the membership table (in Refresh). Local entry: [SiloAddress=S10.244.6.128:11111:150624604 SiloName=Silo_b7daa Status=Dead HostName=memex-portal-deployment-85587f8869-5j9gx ProxyPort=30000 RoleName= UpdateZone=0 FaultZone=0 StartTime=2026-10-10 08:10:04.851 GMT IAmAliveTime=2026-10-10 08:29:36.513 GMT Suspecters=[S10.244.17.58:11111:150624528, S10.244.5.66:11111:150623330, S10.244.17.58:11111:150624528] SuspectTimes=[2026-10-10 08:29:10.921 GMT, 2026-10-10 08:29:10.921 GMT, 2026-10-10 08:29:10.921 GMT]].
FATAL EXCEPTION from Orleans.Runtime.MembershipService.MembershipTableManager. Context: I have been told I am dead, so this silo will stop! Reason: I should be Dead according to the membership table (in Refresh). Local entry: [SiloAddress=S10.244.6.128:11111:150624604 SiloName=Silo_b7daa Status=Dead HostName=memex-portal-deployment-85587f8869-5j9gx ProxyPort=30000 RoleName= UpdateZone=0 FaultZone=0 StartTime=2026-10-10 08:10:04.851 GMT IAmAliveTime=2026-10-10 08:29:36.513 GMT Suspecters=[S10.244.17.58:11111:150624528, S10.244.5.66:11111:150623330, S10.244.17.58:11111:150624528] SuspectTimes=[2026-10-10 08:29:10.921 GMT, 2026-10-10 08:29:10.921 GMT, 2026-10-10 08:29:10.921 GMT]].. Exception: null.\nCurrent stack:    at System.Environment.get_StackTrace()
   at Orleans.Runtime.FatalErrorHandler.OnFatalException(Object sender, String context, Exception exception) in /_/src/Orleans.Runtime/Core/FatalErrorHandler.cs:line 32
   at Orleans.Runtime.MembershipService.MembershipTableManager.KillMyselfLocally(String reason) in /_/src/Orleans.Runtime/MembershipService/MembershipTableManager.cs:line 708
   at Orleans.Runtime.MembershipService.MembershipTableManager.CheckIfLocalSiloIsDead(String caller) in /_/src/Orleans.Runtime/MembershipService/Membershi…[truncated]

Opened automatically from Admin/_LogIncident/0803784e63467766. Recurrences are folded into this issue rather than opening new ones.
It also stands for the whole log site dfcbe511c74e943b: other fingerprints of this site fold in here as comments rather than opening tickets of their own.

Dominant language
C#
Stars
12
Forks
5
Avg merge
4h 15m
Merged PRs (30d)
969

Getting set up

  • No Dockerfile or Docker Compose file
  • Has a pull request template
  • No contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Systemorph/MeshWeaver

All issues in Systemorph/MeshWeaver

Similar issues

More C# issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.