Explicit primary FAILOVER uses a stale replication offset after changing roles
Maintainers usually reply within 2 days
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 55/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- csharp
- Domain
- distributed-systems
Research direction
Start with libs/cluster/Server/Failover/PrimaryFailoverSession.cs and trace how TryStopWrites changes roles before the synchronization target is read from ReplicationManager.ReplicationOffset; consult ReplicationManager.cs and ClusterManager.cs for the related flow. Run the primary-side FAILOVER TO regression with one and two append-only sublogs; done means the target is promoted, retains data, and accepts writes.
Written by the indexing model from the issue text.
Description
Problem
An explicitly requested primary-side FAILOVER TO can return OK without promoting the target replica because the synchronization step reads a stale replication offset.
TryStopWrites changes the local role from primary to replica. The subsequent read of ReplicationManager.ReplicationOffset therefore returns the replica offset instead of the paused primary's final append-only log address.
Reproduction
- Create a primary and replica with append-only logging enabled.
- Assign slots to the primary, replicate existing data, and wait until the replica has synchronized both data and slot assignments.
- Send this command to the primary:
FAILOVER TO <replica-address> <replica-port> TIMEOUT 30000
- Wait for the target replica to report the primary role.
The command returns OK, but the replica does not become primary. Regression tests timed out after 60 seconds with both one and two append-only sublogs.
Expected behavior
The primary should use a stable synchronization target captured from its append-only log after writes have drained. The synchronized replica should be promoted, retain existing data, and accept new writes.
Scope
This affects the standalone FAILOVER operation explicitly sent to a primary. It is not an automatic failover path and does not apply to a correctly formed replica-side CLUSTER FAILOVER command.
Evidence
The regression tests still timed out after fixing only the separate GarnetClient option-framing defect. Applying the stable-offset fix made them pass, confirming that this is a distinct defect.
Affected files:
libs/cluster/Server/Failover/PrimaryFailoverSession.cslibs/cluster/Server/Replication/ReplicationManager.cslibs/cluster/Server/ClusterManager.cs
- Dominant language
- C#
- Stars
- 12k
- Forks
- 709
- Avg merge
- 2d 9h
- Merged PRs (30d)
- 66
Getting set up
- Ships a Dockerfile or Docker Compose file
- No pull request template
- No contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/garnet
-
GarnetClient.Info sends framed section arguments and fails for non-default sectionsPossibly taken @vazois claimed this 1 day ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Maintainers usually reply within 2 days
-
GarnetClient.Failover sends framed options instead of raw FORCE and TAKEOVER argumentsPossibly taken @vazois claimed this 1 day ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Maintainers usually reply within 2 days
-
[Tests] Vector-set tests wait for replica AOF sync while VSIM is still runningPossibly taken @kevin-montrose claimed this 1 day ago. Open
microsoft/garnet#2223 · 1 assignee ·
Maintainers usually reply within 2 days
-
A log scan whose end lies inside a disk-resident record returns that record truncatedPossibly taken @TedHartMS claimed this 1 day ago. Open
microsoft/garnet#2219 · 1 assignee ·
Maintainers usually reply within 2 days
-
KEYS / SCAN / DBSIZE inside MULTI after a write hang forever, pin a core and leave the key lockedPossibly taken @kevin-montrose claimed this 1 day ago. Open
Difficulty 4/5 3-5 days Newbie friendliness 48/100
microsoft/garnet#2209 · 1 assignee ·
Maintainers usually reply within 2 days
All issues in microsoft/garnet
Similar issues
-
[Tool] DirectBenchOpenhas-image has-readme needs-attention new-tool repo-verified
Difficulty 1/5 1-3 hours Newbie friendliness 62/100
shanselman/TinyToolTown#844 · 2 comments ·
Maintainers usually reply within 3 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
PCL-Community/PCL-CE#3658 ·
Maintainers usually reply within 1 day
-
Deploy & Patch-issues opprettes ikke: create-pnd-issues.yml har feilet hver uke siden 2025-09-08Open
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
Altinn/altinn-auth#4359 ·
Maintainers usually reply within 1 day
-
アプリ: チャット 優先: 中 提案
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
yksr-melt/Meltype#243 · 1 comment ·
Maintainers usually reply within 1 day
-
bug core
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day