Add retry logic to updateRecoveryWindow to handle concurrent ObjectStore status updates
メンテナーはふだん 1 日以内に返信
評価
- 難易度
- 3/5
- 見積もり時間
- 1〜2日
- 初心者へのやさしさ
- 62/100
- issue の種類
- バグ
- 明瞭さ
- 明確に書かれている
- 活発さ
- 停滞
- 技術スタック
- go, kubernetes
調査の方向性
internal/cnpgi/instance/recovery_window.go から始め、updateRecoveryWindow を setLastFailedBackupTime と併せて読み、その後 backup.go と retention.go にある呼び出し元を調べます。変更によって、issue に記載された競合エラーなしに ObjectStore のステータス更新の同時実行を処理しつつ、リカバリウィンドウの更新が維持されることを確認します。
索引モデルが issue の本文から書いたものです。
説明
Problem
When running scheduled backups with retention policies, we observe transient errors:
{"level":"error","msg":"Error while updating the recovery window in the ObjectStore status stanza. Skipping.","error":"Operation cannot be fulfilled on objectstores.barmancloud.cnpg.io \"cluster-name-backup\": the object has been modified; please apply your changes to the latest version and try again"}
{"level":"error","msg":"Retention policy enforcement failed","error":"Operation cannot be fulfilled on objectstores.barmancloud.cnpg.io \"cluster-name-backup\": the object has been modified; please apply your changes to the latest version and try again"}
Root Cause Analysis
After investigating the plugin source code, we identified that the updateRecoveryWindow function in internal/cnpgi/instance/recovery_window.go performs a direct status update without retry logic:
// recovery_window.go:40
func updateRecoveryWindow(...) error {
// ... builds status ...
return c.Status().Update(ctx, objectStore) // No retry on conflict
}
This function is called from two places that can run concurrently:
- backup.go:169 - After a backup completes successfully
- retention.go:66 - During periodic retention policy enforcement (default every 5 minutes)
When both operations happen close together, Kubernetes optimistic concurrency control rejects one update because the resourceVersion changed between read and write.
Evidence
The same file already has a function that correctly handles this scenario:
// recovery_window.go:65 - setLastFailedBackupTime
func setLastFailedBackupTime(...) error {
return retry.RetryOnConflict(retry.DefaultBackoff, func() error {
var objectStore barmancloudv1.ObjectStore
if err := c.Get(ctx, objectStoreKey, &objectStore); err != nil {
return err
}
// ... update status ...
return c.Status().Update(ctx, &objectStore)
})
}
The setLastFailedBackupTime function uses retry.RetryOnConflict which:
- Gets a fresh copy of the resource before updating
- Retries on conflict with exponential backoff
Impact
- Severity: Low - backups complete successfully, status eventually updates
- User experience: Confusing error messages in logs
- Frequency: Depends on backup/retention timing overlap (we see ~2 errors per 24h)
Proposed Fix
Apply the same retry pattern to updateRecoveryWindow:
func updateRecoveryWindow(
ctx context.Context,
c client.Client,
backupList *catalog.Catalog,
objectStore *barmancloudv1.ObjectStore,
serverName string,
) error {
return retry.RetryOnConflict(retry.DefaultBackoff, func() error {
// Get fresh copy
var freshObjectStore barmancloudv1.ObjectStore
if err := c.Get(ctx, client.ObjectKeyFromObject(objectStore), &freshObjectStore); err != nil {
return err
}
// Build recovery window
convertTime := func(t *time.Time) *metav1.Time {
if t == nil {
return nil
}
return ptr.To(metav1.NewTime(*t))
}
recoveryWindow := freshObjectStore.Status.ServerRecoveryWindow[serverName]
recoveryWindow.FirstRecoverabilityPoint = convertTime(backupList.GetFirstRecoverabilityPoint())
recoveryWindow.LastSuccessfulBackupTime = convertTime(backupList.GetLastSuccessfulBackupTime())
if freshObjectStore.Status.ServerRecoveryWindow == nil {
freshObjectStore.Status.ServerRecoveryWindow = make(map[string]barmancloudv1.RecoveryWindow)
}
freshObjectStore.Status.ServerRecoveryWindow[serverName] = recoveryWindow
return c.Status().Update(ctx, &freshObjectStore)
})
}
Environment
- Plugin version: 0.10.0
- CNPG Operator: 1.26+
- Kubernetes: 1.29+
- Object storage: AWS S3
We're happy to submit a PR if this approach looks correct.
- 主要言語
- Go
- スター
- 198
- フォーク
- 81
- 平均マージ
- 2日 6時間
- マージ済み PR(30日)
- 19
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートなし
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
cloudnative-pg/plugin-barman-cloud のほかの issue
-
Wrong `ObjectStore` key logged when the backup object store can't be retrieved in the instance sidecar対応中かも @SamarthSRao が 2 日前に担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
cloudnative-pg/plugin-barman-cloud#1142 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
cloudnative-pg/plugin-barman-cloud#1117 · リアクション 1 件 ·
メンテナーはふだん 1 日以内に返信
-
data.compression rejects zstd although barman-cloud-backup supports it対応中かも @dgsardina が 5 日前に担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
cloudnative-pg/plugin-barman-cloud#1104 · リアクション 4 件 ·
メンテナーはふだん 1 日以内に返信
-
Add RBAC aggregation labels to ObjectStore editor/viewer ClusterRoles対応中かも @stefanpeknik が 23 日前に担当しました。 オープン
難易度 1/5 1時間未満 初心者へのやさしさ 82/100
cloudnative-pg/plugin-barman-cloud#1102 ·
メンテナーはふだん 1 日以内に返信
-
Prefetched WAL segments are not returned until the whole batch finishes対応中かも @Kobayashi-UwU が 1 日前に担当しました。 オープン
難易度 4/5 3〜5日 初心者へのやさしさ 15/100
cloudnative-pg/plugin-barman-cloud#1149 ·
メンテナーはふだん 1 日以内に返信
cloudnative-pg/plugin-barman-cloud の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
メンテナーはふだん 1 日以内に返信
-
duplication
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
openvibely/openvibely#1443 ·
メンテナーはふだん 2 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 60/100
canonical/service-mesh#845 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 62/100
-
難易度 2/5 1〜3時間 初心者へのやさしさ 80/100
keyxmakerx/Chronicle#1179 ·
メンテナーはふだん 1 日以内に返信