dj.Diagram SVG output is not byte-reproducible: set iteration order leaks into node emission order

Abierto
#1,551 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
3/5
Tiempo estimado
1-2 días
Aptitud para principiantes
78/100
Tipo de issue
Error
Claridad
Bien especificado
Estado de actividad
Activo
Stack tecnológico
python

Línea de trabajo

Empieza en diagram.py, en las rutas que consumen conjuntos alrededor de las líneas 1210 y 1251-1252; después, inspecciona los demás bucles de emisión descritos en el issue. Reproduce la generación de SVG en procesos separados con distintos valores de PYTHONHASHSEED y localiza las pruebas del diagrama. Se considera terminado cuando los bytes SVG son idénticos con todas las semillas hash y se conserva el diagrama renderizado.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

bug

dj.Diagram's SVG output is not byte-reproducible across processes. Two identical runs in one environment — same DataJoint, same pydot, same graphviz, same database — produce SVGs that differ in the order nodes and clusters are emitted. The rendered picture is unaffected: layout coordinates, colors, shapes and edges are all identical. Only the emission order moves.

Reproduction

Rendering a single-schema diagram twice in separate processes:

$ python gen.py && cp out.svg a.svg
$ python gen.py && cp out.svg b.svg
$ diff a.svg b.svg | grep -c '^[<>]'
88

The diff is entirely reordering — clust2/clust4 swap, node1/node2 swap, and the <path> coordinates travel with them unchanged:

-<g id="clust2" class="cluster">
-<title>cluster_cluster_entity_Fluorescence</title>
+<g id="clust4" class="cluster">
+<title>cluster_cluster_entity_Segmentation</title>

Pinning the hash seed makes it go away, which identifies the mechanism:

result
two runs, default (randomized) seed differ
two runs, PYTHONHASHSEED=0 identical
PYTHONHASHSEED=0 vs PYTHONHASHSEED=12345 differ

Cause

nodes_to_show and _expanded_nodes are Python sets of table-name strings (diagram.py:245, :261, and the result.nodes_to_show = ... assignments at :470, :489, :556). Several emission-path loops iterate them directly, or iterate the result of a set intersection:

# diagram.py:1210 — dimension marking
valid_nodes = self.nodes_to_show.intersection(set(self.nodes()))
for name in valid_nodes:
    ...

# diagram.py:1251-1252 — collapse
valid_nodes = self.nodes_to_show.intersection(set(self.nodes()))
valid_expanded = self._expanded_nodes.intersection(set(self.nodes()))

str.__hash__ is salted per process by default (PEP 456), so set iteration order for table names varies between runs, and that order reaches the emitted document.

Why it matters

It defeats byte-comparison of generated figures. datajoint-docs commits three generated SVGs and ships scripts/gen_pipeline_diagrams.py --check to detect renderer drift in CI (datajoint/datajoint-docs#266). That check cannot be relied on today: one of the three figures fails against its own committed output roughly half the time for reasons unrelated to the renderer, so a genuine notation change is indistinguishable from reshuffling. It also means every regeneration produces large spurious diffs — the real change gets buried.

More generally, anyone committing a dj.Diagram SVG to version control gets churn on every re-render.

Suggested fix

Sort where the set is consumed, not where it is built — the sets are the right structure for the membership tests they exist for. sorted(...) at the emission-path call sites above is a small, behavior-preserving change; the ordering only needs to be stable, not meaningful. Worth a test that renders the same diagram in two subprocesses with different PYTHONHASHSEED values and asserts the output matches, since an in-process test cannot catch this.

Found while verifying the committed docs figures against released 2.3.3 for datajoint/datajoint-docs#266.

Lenguaje dominante
Python
Estrellas
197
Forks
98
Merge medio
6 d 10 h
PR fusionados (30 d)
4

Guía de contribución

Abrir la guía de contribución

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de datajoint/datajoint-python

Todos los issues de datajoint/datajoint-python

Issues similares

Más issues de Python

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.