Backup timeout: parent and child share one deadline — child cleanup is dead code and the staging dir leaks permanently
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Newbie friendliness
- 68/100
- Issue type
- Bug
- Clarity
- Clearly specified
- Activity status
- Quiet
- Domain
- infrastructure
Research direction
Start in backup_daily.py at lines 1076-1079, 1050, and 371, then reproduce the parent timeout with a child using TemporaryDirectory under output_dir. Add regression coverage that forces the parent deadline and verifies no brainlayer-backup-* directory remains; done means child cleanup runs and its receipt is written before the parent intervenes.
Written by the indexing model from the issue text.
Description
Found by reviewer-a2 during the exact-tip review of #683 @ 8c678c04 (verdict ACCEPT — this is a follow-up, not a blocker). Reproduced by running it, not by reading it.
The defect
main() resolves _configured_backup_timeout_seconds() once and hands the same T to both sides:
- parent arms
child.wait(timeout=T) - the re-exec'd child arms
setitimer(ITIMER_REAL, T)off the same inherited env var
The parent starts its clock first. Child interpreter + import measured at 0.07s, so the parent's deadline always expires first. The child therefore never reaches its except BackupTimeoutError path — it takes SIGTERM with default disposition: immediate death, no finally, no context-manager unwind.
backup_daily.py:1076-1079, :1050, :371.
Reproduced
Supervising a child holding TemporaryDirectory(prefix="brainlayer-backup-", dir=output_dir) returned 124 and left brainlayer-backup-nm9x2xug/snapshot.db on disk.
Nothing in the repo ever sweeps that prefix. The only glob is .*.db.attempt-* (:242), which does not match it.
Why it matters — this is the lane's 4th failure wearing a new hat
A 6-hour timeout fires while gzipping the 13.9 GB snapshot. The tmpdir holds a full-size raw snapshot plus a partial .gz, and leaks permanently.
Two such nights and required_bytes = db_size*3 + 512MB + reserve (:359) exceeds free space — backups then fail on the space check. The lane that exists to fix a 45-day silent backup outage would produce a new silent backup outage, by a different route.
Fix — one line
Give the parent T + grace (e.g. +300) so the child reaches its own cleanup and writes its own receipt, and the parent only intervenes when the child is genuinely wedged — which is the case #683 exists to handle. As written, the parent pre-empts the child every time, so the child's entire timeout path is unreachable code.
Secondary, from the same review
- The deadline is still hardcoded — third constant in this lane.
7200 → 21600(backup_daily.py:56,scripts/launchd/backup-daily.sh:8). Stated plainly: a longer fuse, not a fix. A derived deadline (from DB size and measured throughput) is the actual answer. outside_python_signal_deliverycannot fail against pre-#683 code —_supervise_backup_processdid not exist there (verified atd0e8c746:992). It is a regression guard, not proof of the fix.
RED
Supervise a child that holds a TemporaryDirectory under output_dir, force the parent deadline to fire, and assert no brainlayer-backup-* directory survives. It does today, and nothing sweeps it.
- Dominant language
- Python
- Stars
- 9
- Forks
- 7
- Avg merge
- 2h 8m
- Merged PRs (30d)
- 225
Getting set up
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from EtanHey/brainlayer
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
EtanHey/brainlayer#999 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
EtanHey/brainlayer#986 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
EtanHey/brainlayer#985 ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 88/100
EtanHey/brainlayer#676 ·
Maintainers usually reply within 1 day
-
enhancement
Difficulty 1/5 1-3 hours Newbie friendliness 88/100
EtanHey/brainlayer#612 ·
Maintainers usually reply within 1 day
All issues in EtanHey/brainlayer
Similar issues
-
Difficulty 1/5 Under an hour Newbie friendliness 72/100
letsencrypt/cp-cps#353 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
PedestrianDynamics/pyFDS-Evac#394 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
DOI-USGS/pywatershed#421 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
python-pillow/Pillow#10087 · 1 comment ·
Maintainers usually reply within 1 day