`_generate_sample_rand` seeds a Mersenne Twister per Transaction, eagerly, even when unsampled (6.4 µs)
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 64/100
- Tipo de issue
- Refactorización
- Claridad
- Bastante claro
- Estado de actividad
- Activo
- Stack tecnológico
- python
- Área
- performance
Línea de trabajo
Comienza en sentry_sdk/tracing.py, en Transaction.init, y en sentry_sdk/tracing_utils.py, en _generate_sample_rand; inspecciona cómo se lee _sample_rand y si se requiere compatibilidad de salida. Ejecuta el benchmark del issue para comparar las alternativas y, después, verifica que las transacciones no muestreadas eviten trabajo innecesario mientras se preserva el comportamiento de muestreo determinista.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Summary
Transaction.__init__ unconditionally computes _generate_sample_rand(self.trace_id), and _generate_sample_rand seeds a Mersenne Twister to produce a single float. That is 6.4 µs per call on CPython 3.14 / Apple M2, paid on every request through the ASGI integrations even when tracing is disabled and the value can never be used.
Two independent problems:
1. It is eager. sentry_sdk/tracing.py, Transaction.__init__:
baggage_sample_rand = None if self._baggage is None else self._baggage._sample_rand()
if baggage_sample_rand is not None:
self._sample_rand = baggage_sample_rand
else:
self._sample_rand = _generate_sample_rand(self.trace_id)
_sample_rand is only read when a sampling decision is actually made. With traces_sample_rate unset the transaction is never sampled, so this is pure waste. Making it a lazy property costs nothing.
2. It is expensive. sentry_sdk/tracing_utils.py:
def _generate_sample_rand(trace_id, *, interval=(0.0, 1.0)):
...
rng = Random(trace_id)
sample_rand_scaled = rng.randrange(lower_scaled, upper_scaled)
return sample_rand_scaled / 1_000_000
Random(seed) runs the full MT19937 init_by_array over a 625-word state. Measured with timeit, 50k iterations, best of 5:
| µs | |
|---|---|
Random(trace_id) (32-char hex string) |
6.39 |
Random(int(trace_id, 16)) |
5.90 |
Random(12345) |
5.89 |
int(trace_id, 16) / 2**128 |
0.23 |
The cost is the Mersenne Twister initialisation, not the string hashing - seeding with a small int is just as slow. Deriving a uniformly distributed value in [0, 1) arithmetically from the same trace id is 27x cheaper and just as deterministic.
Impact
On a do-nothing FastAPI endpoint with tracing disabled, making _generate_sample_rand cheap moves the SDK's per-request overhead from +61.3 µs to +53.3 µs (in-process measurement, baseline 15.6 µs/req) - about 13% of the SDK's cost, for a value that is discarded.
Questions
- Is the fix to
Transaction.__init__simply making_sample_randlazy? Happy to open a PR. - Is the exact output of
_generate_sample_randrequired to be bit-compatible across SDKs, or only to be deterministic-from-trace_idand uniformly distributed? If the latter,int(trace_id, 16) / 2**128(scaled into the requested interval) would be a drop-in replacement. If the former, the laziness fix alone still helps.
Repro
import timeit, uuid
from random import Random
tid = uuid.uuid4().hex
n = 50000
for label, fn in [
("Random(hex str)", lambda: Random(tid)),
("Random(int)", lambda: Random(int(tid, 16))),
("Random(12345)", lambda: Random(12345)),
("int(tid,16)/2**128", lambda: int(tid, 16) / 2**128),
]:
print(f"{label:<20} {min(timeit.repeat(fn, number=n, repeat=5)) / n * 1e6:.2f} us")
Environment: CPython 3.14.7, sentry-sdk 2.67.1, Apple M2.
Context: this was found while measuring 7400, where the discarded Transaction is the larger half of the same problem.
- Lenguaje dominante
- Python
- Estrellas
- 2.2k
- Forks
- 672
- Merge medio
- 22 h 47 min
- PR fusionados (30 d)
- 224
Guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de getsentry/sentry-python
-
Python
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
getsentry/sentry-python#7569 · 1 comentario ·
-
Python
Dificultad 1/5 Menos de una hora Aptitud para principiantes 85/100
getsentry/sentry-python#7568 · 2 comentarios ·
-
Python
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
getsentry/sentry-python#7567 · 1 comentario ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
getsentry/sentry-python#7543 · 2 comentarios · 1 asignado ·
-
Python
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
getsentry/sentry-python#6992 · 1 comentario ·
Todos los issues de getsentry/sentry-python
Issues similares
-
essnmx good first issue
Dificultad 1/5 Menos de una hora Aptitud para principiantes 95/100
-
[Feature] 奇物选择添加优先级 Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
syfoud/Simulated_Scepter#174 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
Giskard-AI/giskard-oss#2840 · 1 comentario ·
-
A claim comment carrying the issue number is silently declined while the workflow reports success Abiertoarea: repo bug perceived difficulty: 2
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
yeti-platform/yeti#1380 ·