A simple command line tool to show GPU usage on a SLURM cluster
Repositories
m-bain repositories
Pytorch port of Google Research's VGGish model used for extracting audio features.
🤗 Transformers: State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX.
Extract video features from raw videos using multiple GPUs. We support RAFT and PWC flow frames as well as S3D, I3D, R(2+1)D, VGGish, CLIP, ResNet features.
Implementations of Transformers for Video
Easily create large video dataset from video urls
Large-scale text-video dataset. 10 million captioned short videos.
Robust Speech Recognition via Large-Scale Weak Supervision
OpenAI Whisper ASR Webservice API
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)