Raft upgrade compatibility: member rejoining, WAL page sizes and legacy HTTP headers
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 25/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- csharp
- Domain
- api, databases, distributed-systems
Research direction
Start with the focused upstream HTTP cluster test and the existing downstream membership test for the rejoin failure. Then inspect the WAL fixture and test referenced in PR #403, along with MoveToStandbyState, the Leader setter, UnfreezeAsync, and legacy HTTP header handling. Done means regression coverage passes for member rejoining, old and new WAL page sizes, and absent legacy headers while malformed present values remain rejected.
Written by the indexing model from the issue text.
Description
Two blockers encountered while validating a 6.6.0 → 6.7.2 Raft upgrade
@sakno
Downstream tracking: https://github.com/SlimPlanet/SlimFaas/issues/402. Tested on macOS ARM64 (.NET SDK 10.0.300, 16 KiB system pages), with three real Raft nodes and with focused tests against develop (f10ace6).
1. A removed live member cannot rejoin
Create a three-node HTTP Raft cluster, remove one live follower, append another entry, and add the follower again. Catch-up applies the removal before the node receives the re-addition. Subsequent AppendEntries returns HTTP 500 / QuorumUnreachableException, and the follower does not apply the re-addition.
MoveToStandbyState(resumable: false) faults the election task. The Leader setter subsequently accesses that completed task's Result without checking success. In addition, UnfreezeAsync short-circuits on the previously completed readiness probe, and the old leadership task remains faulted.
The existing downstream membership test passes on 6.6.0 and fails repeatedly on 6.7.2. A focused upstream HTTP cluster test reproduces the failure without downstream application code.
2. Existing WAL metadata pages are interpreted using a different size
WAL metadata pages created with 6.6.0 have a 4096-byte layout. Newer constructors use max(4096, Environment.SystemPageSize), which is 16384 on this host. After correctly restoring the snapshot, reopening the compacted WAL fails with:
WriteAheadLog.InternalException: WAL page 0 doesn't exist on the disk
WriteAheadLog.MetadataPageManager.GetView
WriteAheadLog.ApplyAsync
WriteAheadLog.InitializeAsync
A rolling upgrade fails at its first follower; all 180 synthetic sets were verified on every node before upgrading. The same saved snapshot/WAL restores successfully with 6.6.0. The fixture and test are in https://github.com/SlimPlanet/SlimFaas/pull/403.
The persisted page size needs to survive reopening, including logs already created with the newer larger-page layout. Invalid/mixed page sizes should fail before files are opened or resized. Private-memory buffers also need an alignment compatible with legacy pages smaller than the OS page size.
A proposed fix and tests are being prepared against develop. No causal link to the original downstream staging incident is claimed.
Additional rolling-upgrade blocker: HTTP headers
After correcting WAL restoration, the next real 6.6.0 → 6.7.2 native rolling-upgrade attempt fails because X-Raft-State-Version is required on incoming requests. Legacy nodes have no state version header (implicit version zero). Conversely, a new leader requires X-Raft-Last-Index on AppendEntries responses, which older followers do not emit. Both absent headers need legacy-compatible defaults while keeping malformed present values rejected. Three regression cases fail before this compatibility fix. Tracked with the other fixes in #300.
- Dominant language
- C#
- Stars
- 2k
- Forks
- 159
- Avg merge
- 1d 12h
- Merged PRs (30d)
- 1
Getting set up
Starts the project's dev container in your browser, under your own GitHub account.
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from dotnet/dotNext
-
ai_assisted Lib:Cluster
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
-
Lib:Threading question wontfix
Difficulty 4/5 3-5 days Newbie friendliness 45/100
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
PCL-Community/PCL-CE#3652 ·
Maintainers usually reply within 1 day
-
area:frontend bug FE P3
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
klasolsson81/jobbliggaren#2010 ·
Maintainers usually reply within 1 day
-
agentic-workflows untriaged
Difficulty 1/5 Under an hour Newbie friendliness 65/100
Maintainers usually reply within 1 day
-
area: homeblaze type: bug
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
RicoSuter/Namotion.Interceptor#630 ·
Maintainers usually reply within 1 day
-
Akka.Hosting enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 65/100