Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Catalog maintenance deletes completed Backup objects after a short barman-cloud-backup-list result

オープン
#1,115 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
48/100
issue の種類
バグ
明瞭さ
おおむね明確
活発さ
活発
技術スタック
go, kubernetes, postgresql
領域
cloud, databases

調査の方向性

Read internal/cnpgi/instance/retention.go first, focusing on how GetBackupList results are used during catalog maintenance and how completed Backup objects are selected for deletion. Reproduce or test the short or empty listing case, then verify that transient catalog results no longer remove valid backups while intended retention cleanup still works.

索引モデルが issue の本文から書いたものです。

説明

Versions: plugin-barman-cloud v0.13.0, CloudNativePG 1.30.0, Kubernetes 1.35, S3-compatible object storage (STACKIT).

What happened

On 2026-09-18 the plugin sidecar on the primary of two unrelated clusters logged Deleting backup not in the catalog for the two newest completed Backup objects (the base backups of the 17th and the 18th) and removed them:

{"level":"info","ts":"2026-09-18T09:50:05Z","msg":"Applying backup retention policy","logging_pod":"portal-postgres-2","retentionPolicy":"30d"}
{"level":"info","ts":"2026-09-18T09:50:07Z","msg":"Deleting backup not in the catalog","logging_pod":"portal-postgres-2","backup":"portal-postgres-daily-backup-20260917011500"}
{"level":"info","ts":"2026-09-18T09:50:07Z","msg":"Deleting backup not in the catalog","logging_pod":"portal-postgres-2","backup":"portal-postgres-daily-backup-20260918011500"}

The operator then logged terminal error: Backup.postgresql.cnpg.io "portal-postgres-daily-backup-20260917011500" not found for both objects. The same happened on a second cluster in another namespace at 13:05 UTC.

The bucket itself was intact: ObjectStore.status.serverRecoveryWindow kept its firstRecoverabilityPoint (2026-08-25) and lastSuccessfulBackupTime, WAL archiving never failed, and the next night's scheduled backup completed normally. The kube-apiserver showed a short 5xx spike in the same minute (about 16 errors/s for two five-minute windows, otherwise ~0.02/s), so the listing most likely came back short or empty.

Cause, as far as I can see

internal/cnpgi/instance/retention.go deletes every completed Backup object of the cluster whose status.backupID is not in the current GetBackupList result. There is no guard for the listing itself: an empty or partial catalog (transient object-store or API error, eventual consistency) deletes valid objects. On this side the effect was a false no backup in 26h alert; on a cluster that relies on Backup objects for restore selection it would hide two valid base backups.

Suggestions

  • Skip the deletion pass when the listing is empty, or when it lists fewer backups than the Backup objects that are older than the newest catalog entry.
  • Only delete objects whose backup is older than the retention window, since that is the only case where the catalog is expected to have dropped them.
  • Or require an ID to be missing in two consecutive listings before the object is removed.

Happy to provide more logs or test a fix.

主要言語
Go
スター
192
フォーク
75
平均マージ
1日 4時間
マージ済み PR(30日)
15

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

cloudnative-pg/plugin-barman-cloud のほかの issue

cloudnative-pg/plugin-barman-cloud の issue をすべて見る

似ている issue

Go の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。