Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Trustworthy garbage collection: complete reference discovery (#1469) + race-safe deletion (#1445)

Open
#1,478 5 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
35/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Active
Tech stack
python

Research direction

Start by reading issues #1469 and #1445 to understand the required implementation sequence and the referenced_paths and two-phase deletion requirements. Then review the listed documentation targets, especially reference/specs/garbage-collection.md, how-to/garbage-collection.md, and reference/specs/codec-api.md. Done means the implementation addresses both discovery and race safety, and the documentation covers the codec contract, workflow, configuration, concurrency, and restore semantics.

Written by the indexing model from the issue text.

Description

enhancement

Umbrella tracking the invariant: garbage collection must never delete an object that is still referenced. Today it can, via two independent failure modes that look similar but have different root causes:

  1. Incomplete discovery — #1469: dj.gc.scan hardcodes the built-in codec names (hash/blob/attach, object/npy), so files referenced by a custom codec are never enumerated. The store scan then reports live files as orphans and collect() deletes them. (Row delete also fails to remove custom-codec files.)
  2. Stale discovery / TOCTOU race — #1445: the scan is correct, but an insert during the scan→delete window has its file deleted out from under it.

They are complementary layers of one goal, not duplicates — and must be fixed in order:

  • #1469 — complete reference discovery (correctness, FIRST). Codec-owned referenced_paths hook; scan and delete-cleanup become codec-driven instead of name-hardcoded. Closes active data loss for real pipelines (aeon_mecha).
  • #1445 — race-safe deletion (hardening, SECOND). Two-phase quarantine → grace → purge with re-check-before-delete; _trash/ prefix for state.
Why the sequence matters

#1445's purge() re-check reruns the scan. On a codec-blind scan (pre-#1469) it would still classify a live custom-codec file as an orphan and purge it after the grace window — so the grace window can't save you from a scan that never sees the reference. Complete discovery (#1469) is the precondition for safe deletion (#1445).

Documentation (datajoint-docs) — to land with the implementation

Add

  • reference/specs/garbage-collection.md (new normative spec — none exists today). The orphan-determination model, the referenced_paths codec contract, the two-phase quarantine/grace/purge state machine, config keys (gc.grace_seconds), re-check/concurrency semantics, backend atomic-move requirements, and restore. (#1445 explicitly asks for a written spec.)

Update

  • how-to/garbage-collection.md ("Clean Up Object Storage") — document the two-phase workflow (quarantine / purge / restore, grace_seconds); note custom-codec external files are now handled; revise the "single-pass, best-effort" admonition added in #189 once two-phase lands.
  • reference/specs/codec-api.md — document referenced_paths as part of the Codec contract (required for any codec that owns external artifacts).
  • how-to/create-custom-codec.md, explanation/custom-codecs.md, how-to/use-plugin-codecs.md — author guidance: if your codec writes external files, implement referenced_paths so delete + GC see them (otherwise files leak or, worse, get misclassified as orphans and deleted).
  • reference/specs/provenance.md — the GC concurrency wording references single-pass semantics; align once two-phase ships (minor).
  • Optionally a short explainer (e.g. in explanation/object-storage-overview.md) on how DataJoint tracks external references — codecs own their paths; delete and GC consult them.
Dominant language
Python
Stars
197
Forks
98
Avg merge
6d 7h
Merged PRs (30d)
1

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from datajoint/datajoint-python

All issues in datajoint/datajoint-python

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.