m-bain/whisperX

Potential optimisation: temp-file mode for load_audio() to reduce memory on long audio

Ouverte

#1 440 ouverte le 20 juin 2026

 (2 commentaires) (0 réaction) (0 personne assignée)Python (2 386 forks)batch import
enhancementhelp wanted

Métriques du dépôt

Stars
 (23 718 étoiles)
Métriques de merge PR
 (Métriques PR en attente)

Description

load_audio() currently buffers the entire ffmpeg output in Python memory via subprocess.run(capture_output=True). For very long audio (10h+ ≈ 1.15 GB raw PCM), this can cause OOM on constrained systems since the Python bytes buffer, numpy view, and float32 copy all coexist at peak.

A use_tmp_file=True option could write ffmpeg output to a temp file instead of piping to stdout, then read with np.fromfile(), keeping the raw PCM out of Python's heap entirely.

Originally proposed in #1221 (bundled with unrelated dep changes that have since landed).

Guide contributeur