dj.Diagram SVG output is not byte-reproducible: set iteration order leaks into node emission order

Offen
#1,551 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Dieses Issue hat noch niemand übernommen.

Bewertung

Schwierigkeit
3/5
Geschätzter Aufwand
1-2 Tage
Anfängerfreundlichkeit
78/100
Issue-Typ
Bug
Klarheit
Klar beschrieben
Aktivitätsstatus
Aktiv
Tech-Stack
python

Rechercherichtung

Beginne in diagram.py bei den set-verarbeitenden Pfaden um die Zeilen 1210 und 1251-1252 und untersuche anschließend die anderen im Issue beschriebenen Ausgabeschleifen. Reproduziere die SVG-Erzeugung in separaten Prozessen mit unterschiedlichen PYTHONHASHSEED-Werten und finde die Diagrammtests. Als abgeschlossen gilt die Aufgabe, wenn die SVG-Bytes über alle Hash-Seeds hinweg identisch sind und das gerenderte Diagramm erhalten bleibt.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Beschreibung

bug

dj.Diagram's SVG output is not byte-reproducible across processes. Two identical runs in one environment — same DataJoint, same pydot, same graphviz, same database — produce SVGs that differ in the order nodes and clusters are emitted. The rendered picture is unaffected: layout coordinates, colors, shapes and edges are all identical. Only the emission order moves.

Reproduction

Rendering a single-schema diagram twice in separate processes:

$ python gen.py && cp out.svg a.svg
$ python gen.py && cp out.svg b.svg
$ diff a.svg b.svg | grep -c '^[<>]'
88

The diff is entirely reordering — clust2/clust4 swap, node1/node2 swap, and the <path> coordinates travel with them unchanged:

-<g id="clust2" class="cluster">
-<title>cluster_cluster_entity_Fluorescence</title>
+<g id="clust4" class="cluster">
+<title>cluster_cluster_entity_Segmentation</title>

Pinning the hash seed makes it go away, which identifies the mechanism:

result
two runs, default (randomized) seed differ
two runs, PYTHONHASHSEED=0 identical
PYTHONHASHSEED=0 vs PYTHONHASHSEED=12345 differ

Cause

nodes_to_show and _expanded_nodes are Python sets of table-name strings (diagram.py:245, :261, and the result.nodes_to_show = ... assignments at :470, :489, :556). Several emission-path loops iterate them directly, or iterate the result of a set intersection:

# diagram.py:1210 — dimension marking
valid_nodes = self.nodes_to_show.intersection(set(self.nodes()))
for name in valid_nodes:
    ...

# diagram.py:1251-1252 — collapse
valid_nodes = self.nodes_to_show.intersection(set(self.nodes()))
valid_expanded = self._expanded_nodes.intersection(set(self.nodes()))

str.__hash__ is salted per process by default (PEP 456), so set iteration order for table names varies between runs, and that order reaches the emitted document.

Why it matters

It defeats byte-comparison of generated figures. datajoint-docs commits three generated SVGs and ships scripts/gen_pipeline_diagrams.py --check to detect renderer drift in CI (datajoint/datajoint-docs#266). That check cannot be relied on today: one of the three figures fails against its own committed output roughly half the time for reasons unrelated to the renderer, so a genuine notation change is indistinguishable from reshuffling. It also means every regeneration produces large spurious diffs — the real change gets buried.

More generally, anyone committing a dj.Diagram SVG to version control gets churn on every re-render.

Suggested fix

Sort where the set is consumed, not where it is built — the sets are the right structure for the membership tests they exist for. sorted(...) at the emission-path call sites above is a small, behavior-preserving change; the ordering only needs to be stable, not meaningful. Worth a test that renders the same diagram in two subprocesses with different PYTHONHASHSEED values and asserts the output matches, since an in-process test cannot catch this.

Found while verifying the committed docs figures against released 2.3.3 for datajoint/datajoint-docs#266.

Vorherrschende Sprache
Python
Sterne
197
Forks
98
Ø Merge
6 T. 10 Std.
Gemergte PRs (30 T.)
4

Beitragsleitfaden

Beitragsleitfaden öffnen

Erste Schritte

  1. Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
  3. Forken Sie das Repository und arbeiten Sie in einem Branch.
  4. Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.

Mehr aus datajoint/datajoint-python

Alle Issues in datajoint/datajoint-python

Ähnliche Issues

Weitere Issues zu Python

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.