plugin-barman-cloud does not clean PGDATA before re-extracting the base backup on Job retry, causing disk accumulation across failed attempts
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Newbie friendliness
- 68/100
Research direction
Start at the bootstrap.recovery restore path and the invocation of barman-cloud-restore, focusing on how a Kubernetes Job retry prepares PGDATA. Verify the behavior with repeated failed restore attempts, then ensure each attempt removes PGDATA and adjacent restore artifacts before extraction so disk usage does not accumulate and the original recovery failure remains diagnosable.
Written by the indexing model from the issue text.
Description
Environment
- CloudNativePG operator: 1.26.0
- plugin-barman-cloud: v0.6.0
- PostgreSQL: 17.5
- Restore path:
bootstrap.recovery(externalClusters + barman-cloud plugin), object storage on Azure Blob, restore PVC ~50Gi for a ~8GB source database
What happened
A bootstrap recovery Job failed repeatedly against a backup with a WAL-archiving gap (separate report filed against cloudnative-pg/cloudnative-pg about the Job always retrying the same pinned backup). Kubernetes' default Job.spec.backoffLimit: 6 gave it 7 total pod attempts. The first 6 attempts each:
- Ran
barman-cloud-restoresuccessfully, extracting the full base backup (~8GB) intoPGDATAon the restore PVC. - Started PostgreSQL in recovery, failed WAL replay, and exited.
Nothing cleaned PGDATA between these attempts. Each retry re-extracted the same ~8GB base backup on top of / alongside whatever the previous attempt left behind, so usage accumulated across attempts. By the 7th attempt, the ~50Gi restore PVC was full, and that attempt failed differently and much earlier:
ERROR: Barman cloud restore exception: [Errno 28] No space left on device
Why this is a problem
This turns a clear, correctly-diagnosable failure (a WAL-archiving gap, "WAL ends before end of online backup") into a confusing, unrelated-looking failure (disk exhaustion) purely as a side effect of retrying without cleanup. Whoever investigates the final Job state sees the disk-space error, not the real cause, unless they dig through every earlier retry pod's logs individually. It also means a PVC sized correctly for the database itself is not necessarily sized correctly for backoffLimit + 1 retries of it.
Suggestion
Clean/wipe the target PGDATA directory (and pgdata-adjacent restore artifacts) at the start of each restore attempt, before barman-cloud-restore re-extracts the base backup — or at minimum, detect and fail fast if PGDATA is non-empty going into a retry, rather than silently extracting on top of leftover data from a previous failed attempt.
Happy to provide more logs/detail if useful. This was observed in a production DR setup, not a lab reproduction, so some specifics have been generalized above.
- Dominant language
- Go
- Stars
- 196
- Forks
- 76
- Avg merge
- 6d 8h
- Merged PRs (30d)
- 22
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from cloudnative-pg/plugin-barman-cloud
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
cloudnative-pg/plugin-barman-cloud#1117 · 1 reaction ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
cloudnative-pg/plugin-barman-cloud#1104 · 4 reactions ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 82/100
cloudnative-pg/plugin-barman-cloud#1102 ·
Maintainers usually reply within 1 day
-
Catalog maintenance deletes completed Backup objects after a short barman-cloud-backup-list resultOpen
Difficulty 4/5 3-5 days Newbie friendliness 48/100
cloudnative-pg/plugin-barman-cloud#1115 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
cloudnative-pg/plugin-barman-cloud#1099 · 4 comments ·
Maintainers usually reply within 1 day
All issues in cloudnative-pg/plugin-barman-cloud
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
kind/engineering pulumi/pulumi-terraform Task Workflow Failure
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
pulumi/pulumi-terraform#1215 ·
Maintainers usually reply within 1 day
-
bug needs-acceptance wg/developer-experience-ecosystem
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
vllm-project/semantic-router#4480 ·
Maintainers usually reply within 1 day
-
area/docs kind/documentation priority/backlog triage/accepted
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
lexfrei/cloudflare-tunnel-gateway-controller#943 ·
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
keyxmakerx/Chronicle#967 ·
Maintainers usually reply within 1 day