dj.Diagram SVG output is not byte-reproducible: set iteration order leaks into node emission order

Open
#1,551 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
3/5
Estimated time
1-2 days
Newbie friendliness
78/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python

Research direction

Start in diagram.py at the set-consuming paths around lines 1210 and 1251-1252, then inspect the other emission loops described in the issue. Reproduce the SVG generation in separate processes with different PYTHONHASHSEED values and locate the diagram tests. Done means identical SVG bytes across hash seeds while preserving the rendered diagram.

Written by the indexing model from the issue text.

Description

bug

dj.Diagram's SVG output is not byte-reproducible across processes. Two identical runs in one environment — same DataJoint, same pydot, same graphviz, same database — produce SVGs that differ in the order nodes and clusters are emitted. The rendered picture is unaffected: layout coordinates, colors, shapes and edges are all identical. Only the emission order moves.

Reproduction

Rendering a single-schema diagram twice in separate processes:

$ python gen.py && cp out.svg a.svg
$ python gen.py && cp out.svg b.svg
$ diff a.svg b.svg | grep -c '^[<>]'
88

The diff is entirely reordering — clust2/clust4 swap, node1/node2 swap, and the <path> coordinates travel with them unchanged:

-<g id="clust2" class="cluster">
-<title>cluster_cluster_entity_Fluorescence</title>
+<g id="clust4" class="cluster">
+<title>cluster_cluster_entity_Segmentation</title>

Pinning the hash seed makes it go away, which identifies the mechanism:

result
two runs, default (randomized) seed differ
two runs, PYTHONHASHSEED=0 identical
PYTHONHASHSEED=0 vs PYTHONHASHSEED=12345 differ

Cause

nodes_to_show and _expanded_nodes are Python sets of table-name strings (diagram.py:245, :261, and the result.nodes_to_show = ... assignments at :470, :489, :556). Several emission-path loops iterate them directly, or iterate the result of a set intersection:

# diagram.py:1210 — dimension marking
valid_nodes = self.nodes_to_show.intersection(set(self.nodes()))
for name in valid_nodes:
    ...

# diagram.py:1251-1252 — collapse
valid_nodes = self.nodes_to_show.intersection(set(self.nodes()))
valid_expanded = self._expanded_nodes.intersection(set(self.nodes()))

str.__hash__ is salted per process by default (PEP 456), so set iteration order for table names varies between runs, and that order reaches the emitted document.

Why it matters

It defeats byte-comparison of generated figures. datajoint-docs commits three generated SVGs and ships scripts/gen_pipeline_diagrams.py --check to detect renderer drift in CI (datajoint/datajoint-docs#266). That check cannot be relied on today: one of the three figures fails against its own committed output roughly half the time for reasons unrelated to the renderer, so a genuine notation change is indistinguishable from reshuffling. It also means every regeneration produces large spurious diffs — the real change gets buried.

More generally, anyone committing a dj.Diagram SVG to version control gets churn on every re-render.

Suggested fix

Sort where the set is consumed, not where it is built — the sets are the right structure for the membership tests they exist for. sorted(...) at the emission-path call sites above is a small, behavior-preserving change; the ordering only needs to be stable, not meaningful. Worth a test that renders the same diagram in two subprocesses with different PYTHONHASHSEED values and asserts the output matches, since an in-process test cannot catch this.

Found while verifying the committed docs figures against released 2.3.3 for datajoint/datajoint-docs#266.

Dominant language
Python
Stars
197
Forks
98
Avg merge
6d 10h
Merged PRs (30d)
4

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from datajoint/datajoint-python

All issues in datajoint/datajoint-python

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.