m-bain/whisperX
Potential optimisation: temp-file mode for load_audio() to reduce memory on long audio
Aperta
#1440 aperta il 20 giu 2026
enhancementhelp wanted
Metriche repository
- Star
- (23.680 stelle)
- Metriche merge PR
- (Merge medio 61g 7h) (2 PR mergiate in 30 g)
Descrizione
load_audio() currently buffers the entire ffmpeg output in Python memory via subprocess.run(capture_output=True).
For very long audio (10h+ ≈ 1.15 GB raw PCM), this can cause OOM on constrained systems since the Python bytes buffer, numpy view, and float32 copy all coexist at peak.
A use_tmp_file=True option could write ffmpeg output to a temp file instead of piping to stdout, then read with np.fromfile(), keeping the raw PCM out of Python's heap entirely.
Originally proposed in #1221 (bundled with unrelated dep changes that have since landed).