Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Explicit primary FAILOVER uses a stale replication offset after changing roles

Open
#2,220 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 2 days

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
55/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
csharp

Research direction

Start with libs/cluster/Server/Failover/PrimaryFailoverSession.cs and trace how TryStopWrites changes roles before the synchronization target is read from ReplicationManager.ReplicationOffset; consult ReplicationManager.cs and ClusterManager.cs for the related flow. Run the primary-side FAILOVER TO regression with one and two append-only sublogs; done means the target is promoted, retains data, and accepts writes.

Written by the indexing model from the issue text.

Description

Problem

An explicitly requested primary-side FAILOVER TO can return OK without promoting the target replica because the synchronization step reads a stale replication offset.

TryStopWrites changes the local role from primary to replica. The subsequent read of ReplicationManager.ReplicationOffset therefore returns the replica offset instead of the paused primary's final append-only log address.

Reproduction

  1. Create a primary and replica with append-only logging enabled.
  2. Assign slots to the primary, replicate existing data, and wait until the replica has synchronized both data and slot assignments.
  3. Send this command to the primary:
FAILOVER TO <replica-address> <replica-port> TIMEOUT 30000
  1. Wait for the target replica to report the primary role.

The command returns OK, but the replica does not become primary. Regression tests timed out after 60 seconds with both one and two append-only sublogs.

Expected behavior

The primary should use a stable synchronization target captured from its append-only log after writes have drained. The synchronized replica should be promoted, retain existing data, and accept new writes.

Scope

This affects the standalone FAILOVER operation explicitly sent to a primary. It is not an automatic failover path and does not apply to a correctly formed replica-side CLUSTER FAILOVER command.

Evidence

The regression tests still timed out after fixing only the separate GarnetClient option-framing defect. Applying the stable-offset fix made them pass, confirming that this is a distinct defect.

Affected files:

  • libs/cluster/Server/Failover/PrimaryFailoverSession.cs
  • libs/cluster/Server/Replication/ReplicationManager.cs
  • libs/cluster/Server/ClusterManager.cs
Dominant language
C#
Stars
12k
Forks
709
Avg merge
2d 9h
Merged PRs (30d)
66

Getting set up

  • Ships a Dockerfile or Docker Compose file
  • No pull request template
  • No contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from microsoft/garnet

All issues in microsoft/garnet

Similar issues

More C# issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.