Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

plugin-barman-cloud does not clean PGDATA before re-extracting the base backup on Job retry, causing disk accumulation across failed attempts

Open
#1,100 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
3/5
Estimated time
1-2 days
Newbie friendliness
68/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
azure, go, kubernetes, postgresql
Domain
cloud, databases

Research direction

Start at the bootstrap.recovery restore path and the invocation of barman-cloud-restore, focusing on how a Kubernetes Job retry prepares PGDATA. Verify the behavior with repeated failed restore attempts, then ensure each attempt removes PGDATA and adjacent restore artifacts before extraction so disk usage does not accumulate and the original recovery failure remains diagnosable.

Written by the indexing model from the issue text.

Description

Environment
  • CloudNativePG operator: 1.26.0
  • plugin-barman-cloud: v0.6.0
  • PostgreSQL: 17.5
  • Restore path: bootstrap.recovery (externalClusters + barman-cloud plugin), object storage on Azure Blob, restore PVC ~50Gi for a ~8GB source database
What happened

A bootstrap recovery Job failed repeatedly against a backup with a WAL-archiving gap (separate report filed against cloudnative-pg/cloudnative-pg about the Job always retrying the same pinned backup). Kubernetes' default Job.spec.backoffLimit: 6 gave it 7 total pod attempts. The first 6 attempts each:

  1. Ran barman-cloud-restore successfully, extracting the full base backup (~8GB) into PGDATA on the restore PVC.
  2. Started PostgreSQL in recovery, failed WAL replay, and exited.

Nothing cleaned PGDATA between these attempts. Each retry re-extracted the same ~8GB base backup on top of / alongside whatever the previous attempt left behind, so usage accumulated across attempts. By the 7th attempt, the ~50Gi restore PVC was full, and that attempt failed differently and much earlier:

ERROR: Barman cloud restore exception: [Errno 28] No space left on device
Why this is a problem

This turns a clear, correctly-diagnosable failure (a WAL-archiving gap, "WAL ends before end of online backup") into a confusing, unrelated-looking failure (disk exhaustion) purely as a side effect of retrying without cleanup. Whoever investigates the final Job state sees the disk-space error, not the real cause, unless they dig through every earlier retry pod's logs individually. It also means a PVC sized correctly for the database itself is not necessarily sized correctly for backoffLimit + 1 retries of it.

Suggestion

Clean/wipe the target PGDATA directory (and pgdata-adjacent restore artifacts) at the start of each restore attempt, before barman-cloud-restore re-extracts the base backup — or at minimum, detect and fail fast if PGDATA is non-empty going into a retry, rather than silently extracting on top of leftover data from a previous failed attempt.

Happy to provide more logs/detail if useful. This was observed in a production DR setup, not a lab reproduction, so some specifics have been generalized above.

Dominant language
Go
Stars
196
Forks
76
Avg merge
6d 8h
Merged PRs (30d)
22

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from cloudnative-pg/plugin-barman-cloud

All issues in cloudnative-pg/plugin-barman-cloud

Similar issues

More Go issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.