Add a hidden `_prov` attribute for extrinsic provenance at pipeline boundaries
Personne n'a encore pris cette issue.
Évaluation
- Difficulté
- 5/5
- Temps estimé
- Plus d'une semaine
- Accessibilité débutants
- 35/100
- Type d'issue
- Fonctionnalité
- Clarté
- À clarifier
- Activité
- Active
- Stack technique
- mysql, postgresql, python
- Domaine
- data-engineering, databases
Piste de recherche
Lisez src/datajoint/adapters/base.py ainsi que les chemins des adaptateurs MySQL et PostgreSQL pour comprendre comment les attributs cachés job* sont déclarés et persistés. Examinez ensuite le comportement d’insertion de Entry, Ingest et du fan-out afin de déterminer où un slot _prov pourrait s’appliquer. La finalisation nécessite de résoudre les questions ouvertes et de définir le comportement d’implémentation, de migration, de validation et de propagation.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Description
The gap
Intrinsic provenance is structural: a Compute table's row cannot exist unless its declared upstream exists and is correct, so the foreign-key graph is the lineage. Nothing needs to be recorded for that to hold.
At the boundary, the structure runs out. Rows arrive from outside the pipeline — an Entry table filled by a person or a feed, an Ingest table's make() reading a file the workflow does not track, or a fan-out write into Entry tables that carry no foreign key back to the writer. For those rows the framework has no way to say where they came from, and today each pipeline invents its own: a source_file column here, a notes varchar there, an ingestion log somewhere else, or nothing at all.
That is the one place where DataJoint's provenance story depends entirely on the workflow author's diligence, with no shape to conform to.
The precedent this should follow
The mechanism already exists in the codebase. config.jobs.add_job_metadata adds hidden attributes to Computed/Imported tables at declaration:
_job_start_time datetime(3)
_job_duration float
_job_version varchar(64)
(src/datajoint/adapters/base.py, and per-adapter in mysql.py / the PostgreSQL path.) These are per-row, hidden from heading, written by populate, and durable on the table itself rather than in the job queue. That is the right shape and the right place — it just covers the automated side, where provenance is already intrinsic, and not the boundary, where it is not.
Proposal
A hidden _prov attribute, JSON-typed, on tables that take rows from outside the pipeline.
_prov json DEFAULT NULL # extrinsic provenance for a row that entered from outside
- Where: available on Entry and Ingest tables. A Compute table has no use for it, since its provenance is entailed.
- Who writes it: the workflow author, at the point of entry —
inserton an Entry table, or a fan-out write from inside amake(). The framework provides the slot and the shape; it cannot infer content it did not produce. - What it holds: at minimum the agent, the external source and record identifier, the time, and by what means. A conventional key set matters more than a rigid schema — the point is that two pipelines answering "where did this row come from" answer it in the same shape.
- Master carries it; parts inherit. Consistent with how the master/part relationship works elsewhere.
- Fan-out: inside
make(), the intrinsic record of the ingesting table (its key, its_job_version) is available to propagate into each extrinsic destination, which is what makes a fanned-out row traceable without a foreign key.
Even when the source offers nothing useful — a nightly sync against a colony-management API — the author still records "received from PyRat at 02:15". A slot with a weak value beats no slot.
Why in the framework rather than left to each pipeline
Three things follow from having one shape:
- Export becomes mechanical. W3C PROV wants
wasAttributedToandwasDerivedFromon exactly these rows; OpenLineage wants the same content run-centric. With per-pipeline conventions, every export is bespoke. - The boundary becomes inspectable. "Which Entry rows have no recorded origin" turns into a query rather than an audit.
- ALCOA+ attributability lands where it belongs. Deployments that must answer attributable and contemporaneous for externally-sourced data currently have nowhere standard to put the answer.
Open questions
- Default on or off?
add_job_metadatadefaults toFalse, and tables declared without it never get the columns — a migration edge worth not repeating. A_provslot that is absent on most tables is a slot nobody codes against. - Validation. Enforce a minimal key set at insert, or accept any JSON and let deployments constrain it? Leaning toward the latter in the framework, with strictness as a deployment concern.
- Interaction with
allow_direct_insert. A direct insert into an Ingest table is already a modeling smell (datajoint-docs#267); should it require_prov? - Naming.
_provis short and matches the_job_*convention._sourceor_originwould read more plainly to someone who has not met the term.
Related: datajoint-docs#267 (tier names on the entry/ingest axis), and the fan-out ingestion explanation, which currently tells authors to record source identity without giving them a place to record it.
- Langage dominant
- Python
- Étoiles
- 197
- Forks
- 98
- Merge moyen
- 1 j 23 h
- PR mergées (30 j)
- 6
Préparer son environnement
- Fournit un Dockerfile ou un fichier Docker Compose
- Aucun modèle de pull request
- Lire le guide de contribution
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Autres issues de datajoint/datajoint-python
-
Difficulté 2/5 1-3 heures Accessibilité débutants 76/100
datajoint/datajoint-python#1539 · 3 commentaires ·
-
JSON path equality against a non-string value: silently empty on MySQL, raises on PostgreSQLOuvertebug
Difficulté 3/5 1-2 jours Accessibilité débutants 75/100
datajoint/datajoint-python#1564 ·
-
JSON path type annotation is not portable: `data.n:int` works on PostgreSQL, raises on MySQLOuvertebug
Difficulté 4/5 3-5 jours Accessibilité débutants 68/100
datajoint/datajoint-python#1563 ·
-
enhancement
Difficulté 5/5 Plus d'une semaine Accessibilité débutants 35/100
datajoint/datajoint-python#1562 · 2 commentaires ·
-
bug
Difficulté 3/5 1-2 jours Accessibilité débutants 75/100
datajoint/datajoint-python#1561 · 1 commentaire ·
Toutes les issues de datajoint/datajoint-python
Issues similaires
-
camlight produces degenerate target camera and light framesPeut-être pris @kavyabhand l’a pris aujourd’hui. Ouverte
Difficulté 2/5 1-3 heures Accessibilité débutants 72/100
google-deepmind/mujoco_warp#1743 ·
Les mainteneurs répondent en général sous 1 jour
-
netbird: Update to 0.80.0OuvertePackage: Update Request
Difficulté 2/5 1-3 heures Accessibilité débutants 72/100
getsolus/packages#10933 · 1 commentaire ·
Les mainteneurs répondent en général sous 1 jour
-
[work-item] adam-adae-serious-events: pin missing-AESER filter semanticsPeut-être pris @muse-yamaa-bot l’a pris aujourd’hui. Ouvertework-item
Difficulté 1/5 1-3 heures Accessibilité débutants 90/100
Les mainteneurs répondent en général sous 1 jour
-
[Bug]: Non-vision image fallback calls vision_analyze with an empty source for oversized inline images and tells the model the image is corruptPeut-être pris @liuhao1024 l’a pris aujourd’hui. Ouvertecomp/agent P2 tool/vision type/bug
Difficulté 2/5 1-3 heures Accessibilité débutants 88/100
NousResearch/hermes-agent#132605 ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 2/5 1-3 heures Accessibilité débutants 72/100
pymc-labs/pymc-marketing#3102 · 1 commentaire ·
Les mainteneurs répondent en général sous 1 jour