Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Add retry logic to updateRecoveryWindow to handle concurrent ObjectStore status updates

オープン
#758 コメント 1 件 リアクション 1 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

@gabrielmouallem がすでに取り組んでいます。

2026年2月3日 から。

  • #759 @gabrielmouallem による — オープン

評価

難易度
3/5
見積もり時間
1〜2日
初心者へのやさしさ
62/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
停滞
技術スタック
go, kubernetes

調査の方向性

internal/cnpgi/instance/recovery_window.go から始め、updateRecoveryWindow を setLastFailedBackupTime と併せて読み、その後 backup.go と retention.go にある呼び出し元を調べます。変更によって、issue に記載された競合エラーなしに ObjectStore のステータス更新の同時実行を処理しつつ、リカバリウィンドウの更新が維持されることを確認します。

索引モデルが issue の本文から書いたものです。

説明

bug

Problem

When running scheduled backups with retention policies, we observe transient errors:

{"level":"error","msg":"Error while updating the recovery window in the ObjectStore status stanza. Skipping.","error":"Operation cannot be fulfilled on objectstores.barmancloud.cnpg.io \"cluster-name-backup\": the object has been modified; please apply your changes to the latest version and try again"}

{"level":"error","msg":"Retention policy enforcement failed","error":"Operation cannot be fulfilled on objectstores.barmancloud.cnpg.io \"cluster-name-backup\": the object has been modified; please apply your changes to the latest version and try again"}

Root Cause Analysis

After investigating the plugin source code, we identified that the updateRecoveryWindow function in internal/cnpgi/instance/recovery_window.go performs a direct status update without retry logic:

// recovery_window.go:40
func updateRecoveryWindow(...) error {
    // ... builds status ...
    return c.Status().Update(ctx, objectStore)  // No retry on conflict
}

This function is called from two places that can run concurrently:

  1. backup.go:169 - After a backup completes successfully
  2. retention.go:66 - During periodic retention policy enforcement (default every 5 minutes)

When both operations happen close together, Kubernetes optimistic concurrency control rejects one update because the resourceVersion changed between read and write.

Evidence

The same file already has a function that correctly handles this scenario:

// recovery_window.go:65 - setLastFailedBackupTime
func setLastFailedBackupTime(...) error {
    return retry.RetryOnConflict(retry.DefaultBackoff, func() error {
        var objectStore barmancloudv1.ObjectStore
        if err := c.Get(ctx, objectStoreKey, &objectStore); err != nil {
            return err
        }
        // ... update status ...
        return c.Status().Update(ctx, &objectStore)
    })
}

The setLastFailedBackupTime function uses retry.RetryOnConflict which:

  1. Gets a fresh copy of the resource before updating
  2. Retries on conflict with exponential backoff

Impact

  • Severity: Low - backups complete successfully, status eventually updates
  • User experience: Confusing error messages in logs
  • Frequency: Depends on backup/retention timing overlap (we see ~2 errors per 24h)

Proposed Fix

Apply the same retry pattern to updateRecoveryWindow:

func updateRecoveryWindow(
    ctx context.Context,
    c client.Client,
    backupList *catalog.Catalog,
    objectStore *barmancloudv1.ObjectStore,
    serverName string,
) error {
    return retry.RetryOnConflict(retry.DefaultBackoff, func() error {
        // Get fresh copy
        var freshObjectStore barmancloudv1.ObjectStore
        if err := c.Get(ctx, client.ObjectKeyFromObject(objectStore), &freshObjectStore); err != nil {
            return err
        }

        // Build recovery window
        convertTime := func(t *time.Time) *metav1.Time {
            if t == nil {
                return nil
            }
            return ptr.To(metav1.NewTime(*t))
        }

        recoveryWindow := freshObjectStore.Status.ServerRecoveryWindow[serverName]
        recoveryWindow.FirstRecoverabilityPoint = convertTime(backupList.GetFirstRecoverabilityPoint())
        recoveryWindow.LastSuccessfulBackupTime = convertTime(backupList.GetLastSuccessfulBackupTime())

        if freshObjectStore.Status.ServerRecoveryWindow == nil {
            freshObjectStore.Status.ServerRecoveryWindow = make(map[string]barmancloudv1.RecoveryWindow)
        }
        freshObjectStore.Status.ServerRecoveryWindow[serverName] = recoveryWindow

        return c.Status().Update(ctx, &freshObjectStore)
    })
}

Environment

  • Plugin version: 0.10.0
  • CNPG Operator: 1.26+
  • Kubernetes: 1.29+
  • Object storage: AWS S3

We're happy to submit a PR if this approach looks correct.

主要言語
Go
スター
198
フォーク
81
平均マージ
2日 6時間
マージ済み PR(30日)
19

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

cloudnative-pg/plugin-barman-cloud のほかの issue

cloudnative-pg/plugin-barman-cloud の issue をすべて見る

似ている issue

Go の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。