Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Backup timeout: parent and child share one deadline — child cleanup is dead code and the staging dir leaks permanently

Open
#684 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
3/5
Estimated time
1-2 days
Newbie friendliness
68/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Quiet
Tech stack
python, shell

Research direction

Start in backup_daily.py at lines 1076-1079, 1050, and 371, then reproduce the parent timeout with a child using TemporaryDirectory under output_dir. Add regression coverage that forces the parent deadline and verifies no brainlayer-backup-* directory remains; done means child cleanup runs and its receipt is written before the parent intervenes.

Written by the indexing model from the issue text.

Description

Found by reviewer-a2 during the exact-tip review of #683 @ 8c678c04 (verdict ACCEPT — this is a follow-up, not a blocker). Reproduced by running it, not by reading it.

The defect

main() resolves _configured_backup_timeout_seconds() once and hands the same T to both sides:

  • parent arms child.wait(timeout=T)
  • the re-exec'd child arms setitimer(ITIMER_REAL, T) off the same inherited env var

The parent starts its clock first. Child interpreter + import measured at 0.07s, so the parent's deadline always expires first. The child therefore never reaches its except BackupTimeoutError path — it takes SIGTERM with default disposition: immediate death, no finally, no context-manager unwind.

backup_daily.py:1076-1079, :1050, :371.

Reproduced

Supervising a child holding TemporaryDirectory(prefix="brainlayer-backup-", dir=output_dir) returned 124 and left brainlayer-backup-nm9x2xug/snapshot.db on disk.

Nothing in the repo ever sweeps that prefix. The only glob is .*.db.attempt-* (:242), which does not match it.

Why it matters — this is the lane's 4th failure wearing a new hat

A 6-hour timeout fires while gzipping the 13.9 GB snapshot. The tmpdir holds a full-size raw snapshot plus a partial .gz, and leaks permanently.

Two such nights and required_bytes = db_size*3 + 512MB + reserve (:359) exceeds free space — backups then fail on the space check. The lane that exists to fix a 45-day silent backup outage would produce a new silent backup outage, by a different route.

Fix — one line

Give the parent T + grace (e.g. +300) so the child reaches its own cleanup and writes its own receipt, and the parent only intervenes when the child is genuinely wedged — which is the case #683 exists to handle. As written, the parent pre-empts the child every time, so the child's entire timeout path is unreachable code.

Secondary, from the same review

  • The deadline is still hardcoded — third constant in this lane. 7200 → 21600 (backup_daily.py:56, scripts/launchd/backup-daily.sh:8). Stated plainly: a longer fuse, not a fix. A derived deadline (from DB size and measured throughput) is the actual answer.
  • outside_python_signal_delivery cannot fail against pre-#683 code — _supervise_backup_process did not exist there (verified at d0e8c746:992). It is a regression guard, not proof of the fix.

RED

Supervise a child that holds a TemporaryDirectory under output_dir, force the parent deadline to fire, and assert no brainlayer-backup-* directory survives. It does today, and nothing sweeps it.

Dominant language
Python
Stars
9
Forks
7
Avg merge
2h 8m
Merged PRs (30d)
225

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from EtanHey/brainlayer

All issues in EtanHey/brainlayer

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.