BUG: `SeedSequence.spawn` and `np.random.set_bit_generator` are not thread-safe
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 35/100
Línea de trabajo
Start with the named entry points, SeedSequence.spawn and RandomState._initialize_bit_generator, and reproduce both threaded examples from the issue. Trace how spawn keys and the singleton's bit-generator state are accessed during concurrent work. Done means regression coverage demonstrates unique spawned keys and safe behavior when set_bit_generator races with draws, with the RNG design questions resolved.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Describe the issue:
Two independent thread-safety problems in numpy.random, found while reviewing the fix for gh-32059.
1. SeedSequence.spawn hands out duplicate spawn keys
SeedSequence.spawn reads self.n_children_spawned, constructs the children (Python-level work, so the GIL can be released between bytecodes), and only then increments the counter. Two threads spawning from the same SeedSequence can both observe the same starting index and return children with identical spawn keys, i.e. bit generators that produce identical streams. That silently breaks the central guarantee of the spawning API for parallel work.
The BitGenerator lock can't help here: a SeedSequence is an independent object that can be shared by multiple bit generators (or none), so it needs its own synchronization. Note that a per-instance lock interacts with pickling (the auto-generated cdef pickling would try to serialize it).
BitGenerator.spawn and Generator.spawn delegate to SeedSequence.spawn, so they inherit the bug.
Reproduce:
import sys
import threading
import numpy as np
# not required, just makes it reproduce quickly on the GIL-enabled build
sys.setswitchinterval(1e-5)
ss = np.random.SeedSequence(0)
barrier = threading.Barrier(4)
def spawn(out):
barrier.wait()
out.extend(ss.spawn(500))
results = [[] for _ in range(4)]
threads = [threading.Thread(target=spawn, args=(r,)) for r in results]
for t in threads:
t.start()
for t in threads:
t.join()
keys = [c.spawn_key for r in results for c in r]
print(f"spawned {len(keys)} children, {len(set(keys))} unique spawn keys, "
f"n_children_spawned={ss.n_children_spawned}")
Observed (3 out of 3 runs on a GIL build):
spawned 2000 children, 500 unique spawn keys, n_children_spawned=2000
i.e. all four threads returned the same 500 children.
2. np.random.set_bit_generator races with concurrent draws from the singleton
set_bit_generator calls RandomState._initialize_bit_generator with no synchronization. That method rewrites, non-atomically with respect to running draws:
self._bit_generator(dropping the reference to the old one),self._bitgen, the copiedbitgen_tstruct of function pointers plus state pointer,self._aug_state.bit_generator, andself.lockitself.
Distribution fills run with self.lock, nogil: — without the GIL — reading those function/state pointers, so a concurrent swap lets a fill observe a mixed struct (e.g. MT19937's next_double called on a pcg64_state*, which reads/writes far past the end of the smaller struct) or a state pointer into an already-deallocated bit generator.
Taking the lock inside set_bit_generator would not be sufficient: the lock attribute is replaced by the swap, so a thread still blocked on the old lock and a thread acquiring the new lock can be inside the critical section simultaneously. A real fix needs a stable guard owned by RandomState that is never swapped or the API has to explicitly require that replacement happen only while no other thread uses the singleton.
IMO the simplest fix is to detect and raise an error when we detect that set_bit_generator is used on a shared RandomState.
Reproduce:
import threading
import numpy as np
barrier = threading.Barrier(2)
stop = threading.Event()
def draw():
barrier.wait()
while not stop.is_set():
np.random.random(100_000)
def swap():
barrier.wait()
for i in range(500):
bg = np.random.PCG64(i) if i % 2 else np.random.MT19937(i)
np.random.set_bit_generator(bg)
stop.set()
threads = [threading.Thread(target=draw), threading.Thread(target=swap)]
for t in threads:
t.start()
for t in threads:
t.join()
print("survived")
This script segfaults most times I run it on the GIL-enabled build in my testing.
Python and NumPy Versions:
NumPy main, reproduced using Python 3.14.5.
cc @rkern since fixing these issues requires answering some design questions for the RNG module
- Lenguaje dominante
- Python
- Estrellas
- 32.9k
- Forks
- 12.9k
- Merge medio
- 1 d 4 h
- PR fusionados (30 d)
- 243
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de numpy/numpy
-
04 - Documentation
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
numpy/numpy#32817 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
00 - Bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
numpy/numpy#32394 · 2 comentarios ·
Los mantenedores suelen responder en 1 día
-
04 - Documentation sustain-2026
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
numpy/numpy#32272 · 6 comentarios ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
numpy/numpy#31958 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
04 - Documentation
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
numpy/numpy#31579 · 7 comentarios ·
Los mantenedores suelen responder en 1 día
Todos los issues de numpy/numpy
Issues similares
-
tool-calling
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
vllm-project/vllm#59838 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 92/100
raullenchai/Rapid-MLX#4042 ·
Los mantenedores suelen responder en 1 día
-
documentation
Dificultad 1/5 Menos de una hora Aptitud para principiantes 92/100
transitmatters/mbta-slow-zone-bot#70 ·
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
litestar-org/advanced-alchemy#811 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Los mantenedores suelen responder en 1 día