dj.Diagram SVG output is not byte-reproducible: set iteration order leaks into node emission order
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 3/5
- Geschätzter Aufwand
- 1-2 Tage
- Anfängerfreundlichkeit
- 78/100
- Issue-Typ
- Bug
- Klarheit
- Klar beschrieben
- Aktivitätsstatus
- Aktiv
- Tech-Stack
- python
- Bereich
- data-visualization
Rechercherichtung
Beginne in diagram.py bei den set-verarbeitenden Pfaden um die Zeilen 1210 und 1251-1252 und untersuche anschließend die anderen im Issue beschriebenen Ausgabeschleifen. Reproduziere die SVG-Erzeugung in separaten Prozessen mit unterschiedlichen PYTHONHASHSEED-Werten und finde die Diagrammtests. Als abgeschlossen gilt die Aufgabe, wenn die SVG-Bytes über alle Hash-Seeds hinweg identisch sind und das gerenderte Diagramm erhalten bleibt.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
dj.Diagram's SVG output is not byte-reproducible across processes. Two identical runs in one environment — same DataJoint, same pydot, same graphviz, same database — produce SVGs that differ in the order nodes and clusters are emitted. The rendered picture is unaffected: layout coordinates, colors, shapes and edges are all identical. Only the emission order moves.
Reproduction
Rendering a single-schema diagram twice in separate processes:
$ python gen.py && cp out.svg a.svg
$ python gen.py && cp out.svg b.svg
$ diff a.svg b.svg | grep -c '^[<>]'
88
The diff is entirely reordering — clust2/clust4 swap, node1/node2 swap, and the <path> coordinates travel with them unchanged:
-<g id="clust2" class="cluster">
-<title>cluster_cluster_entity_Fluorescence</title>
+<g id="clust4" class="cluster">
+<title>cluster_cluster_entity_Segmentation</title>
Pinning the hash seed makes it go away, which identifies the mechanism:
| result | |
|---|---|
| two runs, default (randomized) seed | differ |
two runs, PYTHONHASHSEED=0 |
identical |
PYTHONHASHSEED=0 vs PYTHONHASHSEED=12345 |
differ |
Cause
nodes_to_show and _expanded_nodes are Python sets of table-name strings (diagram.py:245, :261, and the result.nodes_to_show = ... assignments at :470, :489, :556). Several emission-path loops iterate them directly, or iterate the result of a set intersection:
# diagram.py:1210 — dimension marking
valid_nodes = self.nodes_to_show.intersection(set(self.nodes()))
for name in valid_nodes:
...
# diagram.py:1251-1252 — collapse
valid_nodes = self.nodes_to_show.intersection(set(self.nodes()))
valid_expanded = self._expanded_nodes.intersection(set(self.nodes()))
str.__hash__ is salted per process by default (PEP 456), so set iteration order for table names varies between runs, and that order reaches the emitted document.
Why it matters
It defeats byte-comparison of generated figures. datajoint-docs commits three generated SVGs and ships scripts/gen_pipeline_diagrams.py --check to detect renderer drift in CI (datajoint/datajoint-docs#266). That check cannot be relied on today: one of the three figures fails against its own committed output roughly half the time for reasons unrelated to the renderer, so a genuine notation change is indistinguishable from reshuffling. It also means every regeneration produces large spurious diffs — the real change gets buried.
More generally, anyone committing a dj.Diagram SVG to version control gets churn on every re-render.
Suggested fix
Sort where the set is consumed, not where it is built — the sets are the right structure for the membership tests they exist for. sorted(...) at the emission-path call sites above is a small, behavior-preserving change; the ordering only needs to be stable, not meaningful. Worth a test that renders the same diagram in two subprocesses with different PYTHONHASHSEED values and asserts the output matches, since an in-process test cannot catch this.
Found while verifying the committed docs figures against released 2.3.3 for datajoint/datajoint-docs#266.
- Vorherrschende Sprache
- Python
- Sterne
- 197
- Forks
- 98
- Ø Merge
- 6 T. 10 Std.
- Gemergte PRs (30 T.)
- 4
Beitragsleitfaden
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus datajoint/datajoint-python
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 76/100
datajoint/datajoint-python#1539 · 3 Kommentare ·
-
Schwierigkeit 5/5 Über eine Woche Anfängerfreundlichkeit 35/100
datajoint/datajoint-python#1550 ·
-
Schwierigkeit 5/5 Über eine Woche Anfängerfreundlichkeit 35/100
datajoint/datajoint-python#1547 ·
-
Schwierigkeit 4/5 3-5 Tage Anfängerfreundlichkeit 52/100
datajoint/datajoint-python#1546 · 1 Kommentar ·
-
Make key_source restrict-only: add key_source_restriction, deprecate parent-redefining overrides Offen
Schwierigkeit 5/5 Über eine Woche Anfängerfreundlichkeit 35/100
datajoint/datajoint-python#1523 · 1 Kommentar ·
Alle Issues in datajoint/datajoint-python
Ähnliche Issues
-
documentation help wanted
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 90/100
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 90/100
simonw/sqlite-utils#872 ·
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 88/100
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 82/100
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100