m-bain/whisperX

Potential optimisation: temp-file mode for load_audio() to reduce memory on long audio

Aperta

#1440 aperta il 20 giu 2026

 (2 commenti) (0 reazioni) (0 assegnatari)Python (2387 fork)batch import
enhancementhelp wanted

Metriche repository

Star
 (23.680 stelle)
Metriche merge PR
 (Merge medio 61g 7h) (2 PR mergiate in 30 g)

Descrizione

load_audio() currently buffers the entire ffmpeg output in Python memory via subprocess.run(capture_output=True). For very long audio (10h+ ≈ 1.15 GB raw PCM), this can cause OOM on constrained systems since the Python bytes buffer, numpy view, and float32 copy all coexist at peak.

A use_tmp_file=True option could write ffmpeg output to a temp file instead of piping to stdout, then read with np.fromfile(), keeping the raw PCM out of Python's heap entirely.

Originally proposed in #1221 (bundled with unrelated dep changes that have since landed).

Guida contributor