m-bain/whisperX

Potential optimisation: temp-file mode for load_audio() to reduce memory on long audio

Aberta

#1.440 aberto em 20 de jun. de 2026

 (2 comentários) (0 reação) (0 responsável)Python (2.386 forks)batch import
enhancementhelp wanted

Métricas do repositório

Stars
 (23.718 estrelas)
Métricas de merge de PR
 (Métricas PR pendentes)

Description

load_audio() currently buffers the entire ffmpeg output in Python memory via subprocess.run(capture_output=True). For very long audio (10h+ ≈ 1.15 GB raw PCM), this can cause OOM on constrained systems since the Python bytes buffer, numpy view, and float32 copy all coexist at peak.

A use_tmp_file=True option could write ffmpeg output to a temp file instead of piping to stdout, then read with np.fromfile(), keeping the raw PCM out of Python's heap entirely.

Originally proposed in #1221 (bundled with unrelated dep changes that have since landed).

Guia do colaborador