Hacktoberfest 2026 : les issues que les mainteneurs ont marquées pour octobre, ouvertes et accessibles aux débutants. Parcourir les issues Hacktoberfest

`_generate_sample_rand` seeds a Mersenne Twister per Transaction, eagerly, even when unsampled (6.4 µs)

Ouverte
#7,401 2 commentaires 0 réactions 0 personnes assignées Voir sur GitHub

Les mainteneurs répondent en général sous 1 jour

Personne n'a encore pris cette issue.

Évaluation

Difficulté
4/5
Temps estimé
3-5 jours
Accessibilité débutants
64/100
Type d'issue
Refactorisation
Clarté
Plutôt claire
Activité
Active
Stack technique
python
Domaine
performance

Piste de recherche

Commencez dans sentry_sdk/tracing.py au niveau de Transaction.init et dans sentry_sdk/tracing_utils.py au niveau de _generate_sample_rand ; examinez comment _sample_rand est lu et si la compatibilité de sortie est requise. Exécutez le benchmark de l’issue pour comparer les alternatives, puis vérifiez que les transactions non échantillonnées évitent tout travail inutile, tout en préservant le comportement d’échantillonnage déterministe.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Description

Python Traces
Summary

Transaction.__init__ unconditionally computes _generate_sample_rand(self.trace_id), and _generate_sample_rand seeds a Mersenne Twister to produce a single float. That is 6.4 µs per call on CPython 3.14 / Apple M2, paid on every request through the ASGI integrations even when tracing is disabled and the value can never be used.

Two independent problems:

1. It is eager. sentry_sdk/tracing.py, Transaction.__init__:

baggage_sample_rand = None if self._baggage is None else self._baggage._sample_rand()
if baggage_sample_rand is not None:
    self._sample_rand = baggage_sample_rand
else:
    self._sample_rand = _generate_sample_rand(self.trace_id)

_sample_rand is only read when a sampling decision is actually made. With traces_sample_rate unset the transaction is never sampled, so this is pure waste. Making it a lazy property costs nothing.

2. It is expensive. sentry_sdk/tracing_utils.py:

def _generate_sample_rand(trace_id, *, interval=(0.0, 1.0)):
    ...
    rng = Random(trace_id)
    sample_rand_scaled = rng.randrange(lower_scaled, upper_scaled)
    return sample_rand_scaled / 1_000_000

Random(seed) runs the full MT19937 init_by_array over a 625-word state. Measured with timeit, 50k iterations, best of 5:

µs
Random(trace_id) (32-char hex string) 6.39
Random(int(trace_id, 16)) 5.90
Random(12345) 5.89
int(trace_id, 16) / 2**128 0.23

The cost is the Mersenne Twister initialisation, not the string hashing - seeding with a small int is just as slow. Deriving a uniformly distributed value in [0, 1) arithmetically from the same trace id is 27x cheaper and just as deterministic.

Impact

On a do-nothing FastAPI endpoint with tracing disabled, making _generate_sample_rand cheap moves the SDK's per-request overhead from +61.3 µs to +53.3 µs (in-process measurement, baseline 15.6 µs/req) - about 13% of the SDK's cost, for a value that is discarded.

Questions
  1. Is the fix to Transaction.__init__ simply making _sample_rand lazy? Happy to open a PR.
  2. Is the exact output of _generate_sample_rand required to be bit-compatible across SDKs, or only to be deterministic-from-trace_id and uniformly distributed? If the latter, int(trace_id, 16) / 2**128 (scaled into the requested interval) would be a drop-in replacement. If the former, the laziness fix alone still helps.
Repro
import timeit, uuid
from random import Random
tid = uuid.uuid4().hex
n = 50000
for label, fn in [
    ("Random(hex str)", lambda: Random(tid)),
    ("Random(int)", lambda: Random(int(tid, 16))),
    ("Random(12345)", lambda: Random(12345)),
    ("int(tid,16)/2**128", lambda: int(tid, 16) / 2**128),
]:
    print(f"{label:<20} {min(timeit.repeat(fn, number=n, repeat=5)) / n * 1e6:.2f} us")

Environment: CPython 3.14.7, sentry-sdk 2.67.1, Apple M2.

Context: this was found while measuring 7400, where the discarded Transaction is the larger half of the same problem.

Langage dominant
Python
Étoiles
2.2k
Forks
672
Merge moyen
22 h 8 min
PR mergées (30 j)
228

Préparer son environnement

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Autres issues de getsentry/sentry-python

Toutes les issues de getsentry/sentry-python

Issues similaires

Plus d'issues Python

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.