dj.Diagram SVG output is not byte-reproducible: set iteration order leaks into node emission order
Nobody has claimed this yet.
Assessment
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Newbie friendliness
- 78/100
- Issue type
- Bug
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- python
- Domain
- data-visualization
Research direction
Start in diagram.py at the set-consuming paths around lines 1210 and 1251-1252, then inspect the other emission loops described in the issue. Reproduce the SVG generation in separate processes with different PYTHONHASHSEED values and locate the diagram tests. Done means identical SVG bytes across hash seeds while preserving the rendered diagram.
Written by the indexing model from the issue text.
Description
dj.Diagram's SVG output is not byte-reproducible across processes. Two identical runs in one environment — same DataJoint, same pydot, same graphviz, same database — produce SVGs that differ in the order nodes and clusters are emitted. The rendered picture is unaffected: layout coordinates, colors, shapes and edges are all identical. Only the emission order moves.
Reproduction
Rendering a single-schema diagram twice in separate processes:
$ python gen.py && cp out.svg a.svg
$ python gen.py && cp out.svg b.svg
$ diff a.svg b.svg | grep -c '^[<>]'
88
The diff is entirely reordering — clust2/clust4 swap, node1/node2 swap, and the <path> coordinates travel with them unchanged:
-<g id="clust2" class="cluster">
-<title>cluster_cluster_entity_Fluorescence</title>
+<g id="clust4" class="cluster">
+<title>cluster_cluster_entity_Segmentation</title>
Pinning the hash seed makes it go away, which identifies the mechanism:
| result | |
|---|---|
| two runs, default (randomized) seed | differ |
two runs, PYTHONHASHSEED=0 |
identical |
PYTHONHASHSEED=0 vs PYTHONHASHSEED=12345 |
differ |
Cause
nodes_to_show and _expanded_nodes are Python sets of table-name strings (diagram.py:245, :261, and the result.nodes_to_show = ... assignments at :470, :489, :556). Several emission-path loops iterate them directly, or iterate the result of a set intersection:
# diagram.py:1210 — dimension marking
valid_nodes = self.nodes_to_show.intersection(set(self.nodes()))
for name in valid_nodes:
...
# diagram.py:1251-1252 — collapse
valid_nodes = self.nodes_to_show.intersection(set(self.nodes()))
valid_expanded = self._expanded_nodes.intersection(set(self.nodes()))
str.__hash__ is salted per process by default (PEP 456), so set iteration order for table names varies between runs, and that order reaches the emitted document.
Why it matters
It defeats byte-comparison of generated figures. datajoint-docs commits three generated SVGs and ships scripts/gen_pipeline_diagrams.py --check to detect renderer drift in CI (datajoint/datajoint-docs#266). That check cannot be relied on today: one of the three figures fails against its own committed output roughly half the time for reasons unrelated to the renderer, so a genuine notation change is indistinguishable from reshuffling. It also means every regeneration produces large spurious diffs — the real change gets buried.
More generally, anyone committing a dj.Diagram SVG to version control gets churn on every re-render.
Suggested fix
Sort where the set is consumed, not where it is built — the sets are the right structure for the membership tests they exist for. sorted(...) at the emission-path call sites above is a small, behavior-preserving change; the ordering only needs to be stable, not meaningful. Worth a test that renders the same diagram in two subprocesses with different PYTHONHASHSEED values and asserts the output matches, since an in-process test cannot catch this.
Found while verifying the committed docs figures against released 2.3.3 for datajoint/datajoint-docs#266.
- Dominant language
- Python
- Stars
- 197
- Forks
- 98
- Avg merge
- 6d 10h
- Merged PRs (30d)
- 4
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from datajoint/datajoint-python
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
datajoint/datajoint-python#1539 · 3 comments ·
-
Difficulty 5/5 Over a week Newbie friendliness 35/100
datajoint/datajoint-python#1550 ·
-
Difficulty 5/5 Over a week Newbie friendliness 35/100
datajoint/datajoint-python#1547 ·
-
Difficulty 4/5 3-5 days Newbie friendliness 52/100
datajoint/datajoint-python#1546 · 1 comment ·
-
Make key_source restrict-only: add key_source_restriction, deprecate parent-redefining overrides Open
Difficulty 5/5 Over a week Newbie friendliness 35/100
datajoint/datajoint-python#1523 · 1 comment ·
All issues in datajoint/datajoint-python
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
-
enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100